What is Moonshot AI Kimi K3: Features, Pricing & Benchmarks
Quick Answer: Kimi K3 is Moonshot AI's flagship open-weight AI model, released July 16, 2026. It has roughly 2.8 trillion...
A Chinese open-weight model just went toe to toe with OpenAI's best, and the internet noticed. Kimi K3, from Moonshot AI, is now scoring close enough to GPT-5.6 Sol that the usual "open source is always behind" narrative doesn't hold anymore.
So does Kimi K3 actually beat ChatGPT? The honest answer: it depends what you're doing with it. Below is the full comparison, pricing included, so you can decide in five minutes instead of reading six benchmark PDFs.
Composite intelligence: GPT-5.6 Sol edges ahead, but the gap is under one point on the leading index.
Coding, especially frontend: Kimi K3 wins outright.
Price per token: Kimi K3 is roughly 40 percent cheaper.
Reliability and polish: ChatGPT still leads, K3's hallucination rate rose in independent testing.
Open weights: Only Kimi K3 offers this, expected by July 27, 2026.
Kimi K3 is Moonshot AI's flagship model, a 2.8 trillion parameter open-weight, multimodal reasoning model, the largest open-weight model released to date, built on a mixture-of-experts architecture with a one-million-token context window and an always-on thinking mode. It went live through Kimi's apps and API on July 16, 2026, with full open weights expected by July 27, 2026, according to Moonshot's own API documentation.
Under the hood, K3 runs on Kimi Delta Attention, a hybrid linear attention mechanism, combined with Attention Residuals, and a sparse Mixture of Experts setup called Stable LatentMoE that activates only 16 of 896 experts per token. Moonshot says this gives K3 roughly 2.5 times the scaling efficiency of its predecessor, K2.
Pricing: one million input tokens cost $0.30 with a cache hit and $3.00 without, and one million output tokens, including reasoning, cost $15.00, regardless of context length.
For a deeper breakdown of K3's full feature set, launch variants, and cost examples, see this complete Kimi K3 features and pricing guide.
ChatGPT's current model family is GPT-5.6, which launched on July 9, 2026, after a limited preview that began June 26. OpenAI split this generation into three tiers instead of one model:
Sol, the flagship, for the hardest reasoning and coding work
Terra, a balanced mid-tier for everyday professional use
Luna, a fast, low-cost tier for high-volume tasks
Inside the ChatGPT app itself, access is layered: GPT-5.5 Instant remains the default for everyday chat, while eligible paid plans can access GPT-5.6 Sol through reasoning settings, and Terra and Luna are mainly available through Work, Codex, and the API.
GPT-5.6 Sol has a 1.05 million token context window, 128K max output, and supports text and image input. Standard pricing is $5 per million input tokens and $30 per million output tokens, with a higher $10/$45 rate applying to requests above 272K input tokens.
|
Category |
Kimi K3 |
GPT-5.6 Sol |
|
Parameters |
2.8 trillion (MoE, open-weight) |
Not publicly disclosed |
|
Context window |
1M tokens |
1.05M tokens |
|
Input pricing |
$3 / 1M tokens |
$5 / 1M tokens |
|
Output pricing |
$15 / 1M tokens |
$30 / 1M tokens |
|
Open weights |
Expected by July 27, 2026 |
No, closed |
|
Multimodal input |
Yes |
Yes |
|
Best for |
Coding, cost efficiency |
Broad reasoning, polish |

On overall intelligence, GPT-5.6 Sol is slightly ahead. According to Artificial Analysis's independent model comparison, Kimi K3 scores about 57 on its Intelligence Index, fourth overall, behind Claude Fable 5 at roughly 60 and GPT-5.6 Sol at roughly 59, and just ahead of Claude Opus 4.8. The same source shows the gap is razor thin: GPT-5.6 Sol's score of 59 versus K3's 57 amounts to roughly half a point, smaller than the index's own stated margin of error.
Coding, especially frontend development, is K3's clearest strength. On LMArena's Frontend Code Arena, K3 took the number one spot, a 17 place jump from its predecessor K2.6, and placed first in six of seven frontend categories, ahead of Claude Fable 5. It also holds its own in sustained coding work, leading on SWE Marathon, a test of long coding sessions, at a score of 42.0, and edging past Fable 5 on Terminal Bench 2.1 with a score of 88.3 versus 84.6, though it trails GPT-5.6 Sol there by half a point.
A broader comparison sums it up: Sol leads on six of nine shared benchmarks including DeepSWE and Terminal-Bench, while K3 leads on three, including FrontierSWE, BrowseComp, and Toolathlon, at roughly 40 percent lower cost per token.
Sol's edge shows up in controlled reasoning and broad agentic tasks. On the GDPval v2 agentic benchmark, K3 sits second behind Fable 5, ahead of Opus 4.8, GLM-5.2, and GPT-5.5, meaning OpenAI and Anthropic's top models still lead on complex, multi-domain agent work. Sol also has a higher-effort mode: GPT-5.6 Sol's Ultra setting reaches 91.9 percent on Terminal-Bench 2.1, ahead of anything K3 currently posts.
Reliability is a real concern for K3 right now. Independent testing measured a 50.9 percent hallucination rate for K3 on one factual-reliability benchmark, up from 39.3 percent for its predecessor. That matters if you're using the model for research or fact-heavy work rather than coding. AI commentator Simon Willison's launch-day writeup noted that Moonshot's own self-reported benchmarks showed K3 beating Claude Opus 4.8 and GPT-5.5 while still trailing Claude Fable 5 and GPT-5.6 Sol, a pattern independent testing has since largely confirmed.
This is where Kimi K3 pulls ahead decisively. At standard rates, K3 costs roughly 40 percent less per token than GPT-5.6 Sol for comparable output, and once open weights arrive, teams will be able to self-host the model entirely, cutting costs further for high-volume workloads.
ChatGPT is mostly a subscription product for individual users, running Free through Plus at $20 a month, two Pro tiers at $100 and $200 a month, Go at $8, and Business at $25 per seat, with model access varying by tier and rollout stage. If you want a flat monthly cost bundled with search, memory, voice, and file tools, ChatGPT's subscription is simpler. If you're running high API volume, K3's per-token pricing and pending open weights are hard to beat.
Benchmark method disputes. The author of one coding benchmark publicly noted that Moonshot used a scoring method his team does not recommend, one that can overstate results.
Max-effort settings only. All of Kimi's published K3 benchmark results were achieved at maximum or high thinking intensity, the slowest and most expensive way to run the model, not typical everyday use.
Conversational polish gap. Reviewers note that despite benchmark parity in some areas, K3's subjective user experience still trails Claude Fable 5 and GPT-5.6 Sol, and the model tends to act rather than ask for clarification in ambiguous scenarios.
Weights aren't public yet. As of this writing, K3 is accessible only via API and Kimi's own apps. The full open-weight release, and its exact license terms, are still pending.
Choose Kimi K3 if:
You build frontend interfaces or run long, sustained coding sessions
Cost per token matters, or you plan to self-host once weights drop
You want an open-weight model for full deployment control
Choose ChatGPT (GPT-5.6) if:
You want the top score on broad, composite intelligence benchmarks
You need low-hallucination, polished conversational output
You're already using Work, Codex, Canvas, memory, or deep research
You'd rather pay a flat monthly fee than manage API billing
There's no single winner, and that's the more interesting story here. Kimi K3 proved an open-weight model can sit within a point or two of the best closed models from OpenAI and Anthropic, something that wasn't true a few months ago. GPT-5.6 Sol still leads on broad reasoning, reliability, and agentic consistency, wrapped in a more polished, ecosystem-rich product.
If your work is coding-heavy and cost-sensitive, test Kimi K3 for yourself. For a full rundown of its features, launch variants, and real cost examples, check the detailed Kimi K3 pricing and benchmarks guide. If you need dependable, low-hallucination output across a wide range of tasks and don't mind paying for it, ChatGPT still earns its price tag.
For frontend web development, yes, K3 currently ranks first on the Frontend Code Arena. For broader software engineering benchmarks like DeepSWE, GPT-5.6 Sol still leads.
Kimi K3 is meaningfully cheaper per token, roughly 40 percent less than GPT-5.6 Sol at standard rates, and will get cheaper still for self-hosters once open weights ship.
Yes, at 2.8 trillion parameters K3 is the largest open-weight model released to date. OpenAI has not disclosed GPT-5.6's parameter count.
Not fully. It's available through Moonshot's API and app today. Full open weights are expected by July 27, 2026, and Moonshot has indicated they'll likely follow the Modified MIT license used for its K2 family, though the exact license file for K3 has not been published yet.
Use caution. Independent testing found its hallucination rate rose compared to its predecessor, so it's currently stronger for coding tasks than fact-heavy research.
Quick Answer: Kimi K3 is Moonshot AI's flagship open-weight AI model, released July 16, 2026. It has roughly 2.8 trillion...
Ai Governance
As AI spreads through business and daily life, so do the moments when it goes wrong, and those moments now...