Kimi K3 vs ChatGPT (GPT-5.6): Which AI Actually Wins in 2026?
A Chinese open-weight model just went toe to toe with OpenAI's best, and the internet noticed. Kimi K3, from Moonshot...
Quick Answer: Kimi K3 is Moonshot AI's flagship open-weight AI model, released July 16, 2026. It has roughly 2.8 trillion total parameters, a 1,048,576-token context window, native vision, and always-on reasoning. API pricing is $3 per million input tokens and $15 per million output tokens, with a discounted $0.30 cache-hit rate. Moonshot says it performs competitively with Claude Fable 5 and beats Claude Opus 4.8 and GPT 5.6 Sol on several coding and agentic benchmarks.
Kimi K3 is Moonshot AI's next-generation large language model, launched on July 16, 2026, as the successor to the Kimi K2 family. Moonshot AI, the Beijing-based startup backed by Alibaba, released Kimi K3 as a 2.8 trillion parameter model that the company says is now the largest open source AI model in the world, with benchmarks showing it performing close to the most powerful proprietary systems from Anthropic and OpenAI. The release was timed to land just ahead of the 2026 World Artificial Intelligence Conference in Shanghai.
According to reporting on Moonshot's own comparisons, the model is open weight, meaning its parameters will be available for users to download and customize, and it reportedly outperforms all rivals except Claude Fable 5 and GPT 5.6 on overall capability, per the company's claims.
Note on independent verification: the benchmark and capability claims in this article are what Moonshot AI has published or stated publicly, plus what named outlets and evaluators (Artificial Analysis, Frontend Code Arena, Bank of America analysts) have independently reported. Where a figure is a company claim rather than a third-party test, this article labels it as such.
Official source: Moonshot AI Kimi K3
2.8 trillion total parameters. Moonshot describes Kimi K3 in its technical blog as the world's first open 3T-class system and the largest open-weight AI model to date.
1 million token context window. The model has a 1,048,576-token context window, one of the largest available on any current model, with flat pricing across the full window (no length-based tiering).
Sparse mixture-of-experts design. K3 activates just 16 of its 896 experts per token, roughly 1.8 percent of the total pool, which keeps inference efficient despite the huge parameter count.
Kimi Delta Attention (KDA) and Attention Residuals. The model is built on two internally developed architectural innovations: Kimi Delta Attention, a hybrid linear attention mechanism, and Attention Residuals, a drop-in replacement for residual connections that delivers consistent scaling gains. Both were previously published as open research by the Moonshot team on GitHub.
Native multimodal and vision understanding. K3 features native visual understanding alongside an always-on reasoning mode the company calls "thinking mode."
Two launch variants. K3 Max is built for chat and agent tasks, while K3 Swarm Max is designed for large-scale parallel processing.
Strong agentic and coding orientation. Moonshot calls K3 its most powerful open-source coding model to date, one that can operate with minimal human oversight, sustain long engineering sessions, navigate massive repositories, and orchestrate terminal tools.
OpenAI SDK compatibility. K3's API is compatible with the OpenAI SDK, lowering the integration barrier for developers already building on that ecosystem.
Open-weight license (pending). Moonshot says full weights will ship under a Modified MIT license. See the status box below for the current release state.
Custom compiler tooling. Moonshot also built MiniTriton, a Triton-like compiler, and recommends serving K3 on supernodes of 64 or more accelerators to keep expert-parallel traffic inside one high-bandwidth domain.
As of July 20, 2026: not yet fully downloadable. Kimi K3 launched API-first on July 16, 2026. Full model weights were scheduled for release by July 27, 2026, under a Modified MIT license, according to researchers who reviewed Moonshot's technical documentation. Until those files appear on Hugging Face, K3 is accessible only through the hosted API and the Kimi app, not as a self-hostable download. If you're budgeting around self-hosting to skip API costs, plan around the July 27 date and confirm availability directly before committing.
This is a live fact that can change quickly. Bookmark this section if you're evaluating self-hosting.
The Kimi K3 API went live July 16, 2026, at api.moonshot.ai/v1 with the model ID kimi-k3.
|
Item |
Price |
|
Input tokens (cache miss) |
$3.00 per million tokens |
|
Input tokens (cache hit) |
$0.30 per million tokens (90% discount) |
|
Output tokens |
$15.00 per million tokens |
|
Context pricing |
Flat across the full 1,048,576-token window, no length tiering |
|
Web search calls |
$0.004 per call, billed separately |
|
Consumer app (Adagio) |
Free, unlimited basic chat, file uploads, web search |
|
Consumer paid tiers |
$19 to $199 per month depending on agent/coding usage |
|
Launch promo |
10 to 30 percent bonus credits on API top-ups, July 15 to August 11, 2026 |
How this compares to other frontier models (verified July 20, 2026):
|
Model |
Input ($/M) |
Output ($/M) |
Notes |
|
Kimi K3 |
$3.00 |
$15.00 |
$0.30 on cache hit |
|
Claude Sonnet 5 (now, through Aug 31, 2026) |
$2.00 |
$10.00 |
Introductory rate, temporarily cheaper than K3 |
|
Claude Sonnet 5 (from Sept 1, 2026) |
$3.00 |
$15.00 |
Standard rate, matches K3 exactly |
|
Claude Opus 4.8 |
$5.00 |
$25.00 |
Unchanged since Opus 4.5 |
|
Claude Fable 5 |
$10.00 |
$50.00 |
Mythos-class, most capable widely available model |
|
GPT-5.6 Sol |
$5.00 |
$30.00 |
Surcharge applies above 272K input tokens |
|
Kimi K2.6 (Moonshot's own prior model) |
~$0.95 |
~$4.00 |
Roughly a quarter of K3's rate |

K3's $15 output rate undercuts Claude Opus 4.8 ($25) and GPT-5.6 Sol ($30), and its $3 input rate matches what Claude Sonnet 5 will cost once its introductory pricing expires on September 1, 2026. Right now, though, Sonnet 5 is running a temporary $2/$10 promotional rate, making it the cheaper of the two until that window closes. Within Moonshot's own lineup and the wider Chinese open-model field, K3 is priced at a premium: it costs roughly three to four times more than its own predecessor, Kimi K2.6, and reportedly sits well above cheaper open models like DeepSeek V4 Pro and GLM 5.2, though those specific ratios come from a single pricing tracker and haven't been cross-confirmed against official DeepSeek and Zhipu rate cards.
Worth knowing before you commit: the most common complaint from early users is that K3 uses more tokens than Claude Fable 5 to finish comparable tasks. Combined with always-on reasoning billed at the $15 output rate, your real cost per completed task can run higher than the sticker price implies.
For a rough sense of real-world spend, based on Moonshot's published rate card:
A medium coding task (around 5,000 input tokens of repo context, 3,000 output tokens of code and reasoning) with a cache miss would cost approximately $0.015 input + $0.045 output = about $0.06 per task.
Run that same task 1,000 times a month (a moderate coding-agent workload) and you're looking at roughly $60/month in raw token cost, before accounting for K3's tendency to use more tokens per task than Fable 5.
With prompt caching applied to repeated repo context, that same workload could drop closer to $20 to $30/month.
These are illustrative estimates using Moonshot's official per-token rates, not guaranteed figures. Actual costs depend heavily on prompt length, caching behavior, and reasoning verbosity.
Official pricing source: Kimi K3 on OpenRouter
|
Benchmark / Source |
Result |
Comparison |
|
Frontend Code Arena (blind developer testing) |
1,679 points, ranked #1 |
Ahead of Claude Fable 5 |
|
Overall capability (Moonshot's own evaluation suite) |
Company-claimed top 3 |
Trails Claude Fable 5 and GPT 5.6 Sol; beats Claude Opus 4.8 and GPT 5.5 |
|
Artificial Analysis (independent estimate) |
Roughly Opus 4.8 / GPT 5.5 tier |
Behind Fable 5 on some arena prompts, per early user reports |
|
Coding and general agent tasks (Moonshot claim) |
"Substantially outperformed" |
Claude Opus 4.8, GPT 5.6 Sol, GPT 5.5 |

Important context on these numbers: most of the capability claims above come directly from Moonshot's own release materials and press briefings, echoed by outlets like Fortune, Forbes, and CNBC. The Frontend Code Arena result is the clearest third-party, blind-tested data point currently public. Named academic benchmark suites (SWE-bench, LiveCodeBench, MMLU, GPQA) had not been independently published with numeric K3 scores at the time of this update. Treat company-reported "beats X" claims as directional until independent leaderboards catch up, and check back as more third-party evaluations land.
Analyst reaction: Bank of America analysts praised the release, noting that even with limited access to the most advanced chips, Moonshot demonstrated it can make major advancements by improving how it trains and designs its models. The market reaction was immediate: shares of Chinese competitors Zhipu and MiniMax fell 28.4 percent and 15.6 percent respectively in Hong Kong trading, and Z.ai shares dropped 28.4 percent as well.
Industry perspective: Simon Koser, chief product officer at AI startup Tzafon, said Kimi K3 is legitimately impressive in areas like coding, and that developers at AI labs could find it compelling. Cursor, the AI coding startup, had previously used earlier Kimi models to help build its Composer 2 coding agent, an early adoption signal for the Kimi model family.
Benchmark source: VentureBeat coverage of Kimi K3
Not quite, on Moonshot's own account. The company says K3 performs competitively with Claude Fable 5, currently the most advanced widely available model on the market, but still trails it on overall performance. K3 does reportedly beat the tier below Fable 5, Claude Opus 4.8, on coding and general agent benchmarks.
On price, the picture is close but not identical. K3's $3/$15 rate matches what Claude Sonnet 5 will cost starting September 1, 2026, once its introductory pricing ends. Right now, Sonnet 5 is temporarily cheaper at $2/$10 during its promotional window. Against Claude Opus 4.8 ($5/$25) and Claude Fable 5 ($10/$50), K3 is meaningfully cheaper on paper, though it is not yet a downloadable open-weight model, while Claude offers a proprietary, fully supported API with no self-hosting question mark.
Similar picture. Moonshot says K3 trails GPT 5.6 Sol on overall performance but substantially outperforms GPT 5.5 on coding and agentic benchmarks. On price, K3 undercuts GPT-5.6 Sol on both input ($3 vs $5) and output ($15 vs $30). So the honest positioning is: K3 sits just below the very top tier of closed models (Fable 5, GPT 5.6 Sol) and clearly ahead of the tier just below that (Opus 4.8, GPT 5.5), at a lower price than either GPT-5.6 Sol or Opus 4.8, at least by Moonshot's own reported testing.
This is where price becomes the deciding factor rather than raw capability. At 2.8 trillion parameters, K3 is roughly 75 percent larger than DeepSeek V4 Pro's approximately 1.6 trillion parameters. But DeepSeek V4 Pro is reportedly dramatically cheaper, by one estimate around a sixth of K3's input cost, with GLM 5.2 running at roughly half K3's price. Those specific ratios come from a single pricing tracker rather than official DeepSeek and Zhipu rate cards, so treat them as directionally useful rather than exact. If your priority is capability and you can absorb frontier-level pricing, K3 is the stronger pick on current benchmarks. If cost per token is the deciding factor, DeepSeek V4 Pro and GLM 5.2 are the cheaper matchups worth testing first, ideally after confirming their current official rates.
Coding agent builders who need a large context window and strong autonomous multi-step performance, and can tolerate frontier-tier pricing.
Teams evaluating open-weight alternatives to Claude and GPT who want a model they can eventually self-host once weights land on July 27.
Not ideal for cost-sensitive, high-volume applications. If you're running a customer support bot or high-frequency low-complexity workload, K3's token-hungry behavior and $15 output rate make cheaper models like K2.6, DeepSeek V4 Pro, or GLM 5.2 a better fit.
Not yet ideal for self-hosting requirements. If your project depends on running weights on your own infrastructure today, K3 isn't there yet. Wait for the July 27 weights release before committing.
Master the EU AI Act
Stay ahead of the latest EU AI Act requirements with expert-led, self-paced training. Learn how the regulation impacts businesses, understand AI risk classifications, compliance obligations, governance frameworks, and practical implementation strategies. Earn a recognized PDF certificate at no extra cost. Build the knowledge and confidence to help your organization achieve AI Act compliance and responsibly deploy AI systems.
Enroll Now →Kimi K3's release lands in the middle of an intensifying US-China AI race. The release comes as global businesses increasingly question the cost of deploying models from Anthropic and OpenAI, and analysts have noted that Moonshot achieved this scale despite limited access to the most advanced chips under US export rules. Some of the reaction to K3's launch has itself become politically charged, with debate in Washington over whether US companies should use Chinese open-source models at all. Congress passed a bill in January 2026 to close an offshore cloud rental loophole that had given Chinese firms remote access to restricted accelerators, adding another layer of scrutiny to how models like K3 are trained and served.
Not fully yet. Moonshot says weights will release under a Modified MIT license by July 27, 2026. Until then, it's API-only.
By Moonshot's own reported benchmarks, K3 beats GPT 5.5 on coding and agentic tasks but trails GPT 5.6 Sol on overall performance. It is also cheaper than GPT-5.6 Sol on both input and output tokens.
The Kimi app has a free Adagio tier with unlimited basic chat. The API is pay-as-you-go starting at $3 per million input tokens.
$3.00 per million input tokens, $15.00 per million output tokens, and $0.30 per million tokens on cache hits.
It's cheaper than Claude Opus 4.8 and Claude Fable 5. It's currently more expensive than Claude Sonnet 5's temporary introductory rate ($2/$10), but will match Sonnet 5's standard rate exactly once that promotion ends on September 1, 2026.
1,048,576 tokens, priced flat with no length-based tiering.
Moonshot targeted July 27, 2026, for the full weights release under a Modified MIT license.
Yes, it has native visual understanding built in alongside always-on reasoning.
DeepSeek V4 Pro is reportedly several times cheaper on input tokens, though K3 is a larger, differently positioned model, and these price ratios come from a single pricing source.
High-volume, cost-sensitive applications are likely better served by cheaper models like Kimi K2.6, DeepSeek V4 Pro, or GLM 5.2, since K3 tends to use more tokens per task at a higher output rate.
A Chinese open-weight model just went toe to toe with OpenAI's best, and the internet noticed. Kimi K3, from Moonshot...
Ai Governance
As AI spreads through business and daily life, so do the moments when it goes wrong, and those moments now...