AIGP: Artificial Intelligence Governance Professional Certification
Artificial intelligence is spreading across workplaces faster than many organizations can create the rules needed to control it. The Artificial...
What changed since our last update: Anthropic released Claude Opus 5 on July 24, 2026, replacing Opus 4.8 as Claude's flagship at the same $5/$25 per-million-token price. It now leads the Artificial Analysis Intelligence Index. Kimi K3 launched July 16, 2026, and its open weights are due the same day this post was last verified. Gemini 3.5 Pro remains delayed with no confirmed release date. Claude Sonnet 5's introductory pricing ($2/$10 per 1M tokens) runs through August 31, 2026, then rises to $3/$15.
Three names keep coming up in every "which AI should I use" conversation right now: Kimi K3, Claude, and Gemini. That's not an accident, since each represents a different bet on where AI is headed. Moonshot AI's Kimi K3 is a 2.8-trillion-parameter open-weight model that became the largest open-weight model released to date when it shipped. Claude is Anthropic's closed-weight, reliability-focused lineup, now headed by the newly released Opus 5 alongside Sonnet 5 and the Mythos-tier Fable 5. Gemini is Google's deeply integrated family built around Gemini 3.1 Pro and the newer 3.6 Flash, with the industry's most aggressive context-window strategy.
The short answer: there's no single winner. Claude Opus 5 currently leads independent composite benchmarks after its July 24 launch, Kimi K3 offers the best price-to-performance ratio and leads on agentic coding benchmarks, and Gemini wins on context window ceiling, native research/search grounding, and multimodal reasoning. The rest of this guide breaks down exactly where each one pulls ahead. Every spec, price, and benchmark below is checked against official vendor documentation and independent trackers like Artificial Analysis, and this comparison was last verified on July 27, 2026.
Best overall: Claude Opus 5 narrowly leads the Artificial Analysis Intelligence Index following its July 24 launch, with Claude Fable 5 and Kimi K3 close behind
Best for coding: Kimi K3 for open-weight, long-horizon agentic coding; Claude for production reliability
Best for writing: Claude, with the strongest tone control, structure, and instruction-following
Best for research: Gemini 3.1 Pro, thanks to native Google Search integration and the largest context window
Best free option: Gemini (Flash tier stays free) or Kimi K3 (generous free quota in the Kimi app)
Best value for money: Kimi K3, offering frontier-class scores at mid-tier API pricing, plus open weights
Best for business users: Claude, with enterprise tooling, Claude Cowork, and Microsoft 365 integration
|
Feature |
Kimi K3 |
Claude |
Gemini |
|
Developer |
Moonshot AI |
Anthropic |
Google DeepMind |
|
Release |
July 16, 2026 |
Sonnet 5 (June 2026), Opus 5 (July 24, 2026), Fable 5 (June 2026) |
Gemini 3.1 Pro (Feb 19, 2026), 3.6 Flash (July 21, 2026) |
|
Model family |
K3 (successor to K2/K2.5/K2.6/K2.7) |
Claude 4.x/5, including Haiku 4.5, Sonnet 5, Opus 5, Fable 5/Mythos 5 |
Gemini 3, including Flash-Lite, Flash, 3.1 Pro, 3.6 Flash |
|
Context window |
1 million tokens |
Up to 1 million tokens |
1 million tokens standard; some sources cite up to 2M on 3.1 Pro |
|
Multimodal support |
Native visual understanding |
Vision, computer-use, document analysis |
Text, image, audio, and video natively |
|
Coding ability |
Very strong, agentic/long-horizon |
Very strong, now Opus 5's lead category |
Strong, agentic terminal work on Flash |
|
Writing quality |
Good, less proven long-form |
Excellent, best-in-class |
Very good, strong for structured content |
|
Reasoning |
Strong, always-on thinking mode |
Excellent, with Opus 5's low/medium/high effort dial |
Excellent, multiple thinking levels plus Deep Think |
|
Web search |
Via Kimi app agent tools |
Available in Claude apps |
Deepest, with native Google Search |
|
Speed |
Moderate, uses more output tokens per task |
Fast; Opus 5 offers a 2.5x Fast Mode at 2x price |
Fast, especially on the Flash tier |
|
Free version |
Yes, quota-limited |
Yes, Sonnet 5 default |
Yes, Flash and Flash-Lite tiers |
|
Paid plans |
¥199 entry tier |
Pro $17 to $20/mo, Max $100 to $200/mo |
AI Plus $4.99, Pro $19.99, Ultra $99.99 to $199.99 |
|
API pricing |
$3 input / $15 output per 1M tokens |
$1 to $10 input / $5 to $50 output depending on tier |
$1.50 to $4 input / $7.50 to $18 output |
|
Best suited for |
Open-weight deployments, long-horizon coding |
Professional writing, enterprise workflows, everyday flagship use |
Research, huge documents, Google ecosystem |
All three are legitimate coding models now, but they win in different ways.
Kimi K3 was built for long-horizon coding, knowledge work, and reasoning, and its launch benchmarks lean hard on agentic and repository-scale tasks. It jumped from #18 to #1 on LMArena's Frontend Code Arena within hours of release, overtaking Claude Fable 5.
Claude's coding crown just changed hands internally. Opus 5 is Anthropic's new flagship for this category. It's a thoughtful, proactive model that comes close to Fable 5's frontier intelligence at half the price, and on coding and knowledge-work evaluations like Frontier-Bench and GDPval-AA it sets a new state of the art for Anthropic. Anthropic's own materials note Opus 5 still trails Fable 5 on some of the hardest, longest-running tasks, so Fable 5 remains the pick for multi-day autonomous work.
Gemini, especially 3.5/3.6 Flash, targets agentic terminal work. Gemini 3.5 Flash scores 76.2% on Terminal-bench 2.1 and 55.1% on SWE-Bench Pro, strong marks for autonomous agent tasks at a lower price than the Pro tier.
A key caveat: benchmark harnesses swing results dramatically. Claude Opus 4.8 scored 69.2% on SWE-bench through Anthropic's own scaffold versus 51.9% on Scale AI's standardized SEAL board, a 17.3-point gap from harness choice alone, and Gemini 3.1 Pro shows a similar 26.4-point spread between its verified and standardized scores. Treat any single vendor-reported coding score with caution. It measures a system, not just a model, and equivalent independent figures for Opus 5 are still emerging days after launch.
Winner: Claude (Opus 5) for production reliability and Anthropic's own state-of-the-art coding claims; Kimi K3 for open-weight, budget-conscious agentic coding at scale.
Claude is consistently the model writers reach for. It holds tone and structure better than the other two across long documents, follows nuanced style instructions more reliably, and produces cleaner prose out of the box for blogs, marketing copy, and email drafting.
Gemini is a strong second, particularly for structured long-form content where its huge context window lets it hold an entire style guide or brand brief in memory alongside the draft.
Kimi K3's writing is competent, but its public track record is overwhelmingly coding and agent focused. There's less independent evidence yet for long-form creative or marketing writing quality compared to Claude and Gemini's years of iteration on this front.
Winner: Claude.
Kimi K3 ships with reasoning on by default. Moonshot calls it "thinking mode," and the model uses max thinking effort by default, with low- and high-effort modes planned for later updates.
Claude's newly released Opus 5 leads the composite reasoning field right now. It's narrowly the most intelligent model on the Artificial Analysis Intelligence Index, offering comparable intelligence to Fable 5 at a meaningfully lower cost per task, and it introduces an effort dial that lets users trade reasoning depth for cost and speed on a per-request basis.
Gemini 3.1 Pro is strongest on formal scientific reasoning. It posted 94.3% on GPQA Diamond and a 77.1% result on ARC-AGI-2 at launch, and Google offers a dedicated Deep Think mode for extended multi-step problems.
Winner: Claude Opus 5 for general composite reasoning; Gemini 3.1 Pro for formal scientific reasoning.
This is Gemini's clearest advantage. Native, first-party Google Search integration gives it an edge on current-events grounding and citation freshness that neither Claude nor Kimi fully matches out of the box. Paired with its large context window, Gemini is well suited to synthesizing many long sources into one report.
Claude's strength is careful, conservative synthesis. It tends to hedge appropriately when sources conflict, which matters when accuracy outweighs speed.
Kimi K3 includes web-connected agent tools through the Kimi app, and Moonshot reports narrow K3 wins on BrowseComp and DeepSearchQA, but it has less of a track record for citation discipline than Claude or Gemini.
Winner: Gemini.
Kimi K3 ships with a native 1-million-token context window. Claude's Sonnet 5 and Opus 5 also run at 1M tokens with no long-context pricing premium on Sonnet 5. Gemini has pushed furthest on paper: every current Gemini model supports at least a 1-million-token context window, with some previews of 3.1 Pro claiming up to 2M, though pricing steps up above 200K tokens on the Pro tier. Gemini 3.1 Pro costs $2.00 per 1M input tokens and $12.00 per 1M output tokens for contexts up to 200K, jumping to $4.00 input / $18.00 output above that.
The chart above shows why this matters for budgeting: raw context size aside, output pricing varies enough between these models that a long-document workflow run at scale can cost meaningfully more on one platform than another.
Winner: Gemini for raw ceiling; Claude for predictable pricing at scale on Sonnet 5.
Gemini leads on multimodal reasoning benchmarks. Gemini 3.1 Pro leads across multimodal reasoning with an 83.6% score on MMMU-Pro, and its native audio/video support outpaces both competitors.
Claude is strong on document-style multimodal work: OCR, charts, tables, screenshots, and UI analysis, with computer-use letting it act directly on what it sees.
Kimi K3 has native visual understanding built into its architecture, but detailed independent OCR/chart benchmarks for K3 specifically are still limited this soon after launch.
Winner: Gemini.
Gemini Flash tiers are generally the fastest of the three for everyday tasks. Claude's Opus 5 offers a Fast Mode running around 2.5 times the default speed at twice Opus 5's base price, and Sonnet 5 is a solid balance for daily work. Kimi K3's benchmarks show it using notably more output tokens per task than comparable models, about 1.9 times GPT-5.6 Sol's total tokens on one evaluation, which can mean longer wait times even where per-token pricing is cheaper.
Winner: Gemini for raw speed; Claude for overall interface polish and reliability.
|
Plan type |
Kimi K3 |
Claude |
Gemini |
|
Free plan |
Yes, quota-limited |
Yes, Sonnet 5 default |
Yes, Flash/Flash-Lite tiers |
|
Monthly subscription |
¥199 entry tier |
Pro $17 to $20/mo, Max $100 to $200/mo |
AI Plus $4.99, Pro $19.99, Ultra $99.99 to $199.99 |
|
API pricing |
$3 / $15 per 1M tokens, around $0.30 cached input |
Sonnet 5 $2 to $3/$10 to $15, Opus 5 $5/$25, Fable 5 $10/$50, Haiku 4.5 $1/$5 |
Gemini 3.6 Flash $1.50/$7.50, Gemini 3.1 Pro $2/$12 up to 200K, $4/$18 above |
|
Enterprise pricing |
Not broadly published yet |
Custom (Team $20 to $125/seat, Enterprise $20/seat plus usage) |
Custom via Workspace/Vertex AI |
|
Best value |
Frontier scores at mid-tier pricing, open weights |
Opus 5 approaches Fable 5 performance at roughly half Fable's price |
Flash tiers at very low per-token cost |
Two notes worth flagging: Claude Sonnet 5 is priced at $2/$10 per million input/output tokens until August 31, 2026, then $3/$15 after that, so locking in usage now is cheaper. And Google's Gemini 3.5 Pro, despite being unveiled in May, still isn't generally available and remains behind Google's internal targets, so treat any 3.5 Pro pricing found online as unconfirmed. Gemini 3.1 Pro remains the actual flagship in production today.
A note on methodology: benchmark harness choice can swing scores by 10 to 26 points for the same model on the same test. A score labeled "Kimi K3" or "Claude Opus 5" always really means "this model plus this specific harness and settings." Use the table below for direction, not decimal-point precision.
|
Benchmark |
Kimi K3 |
Claude |
Gemini |
|
Artificial Analysis Intelligence Index |
57 (fourth overall) |
Opus 5 leads the public snapshot at 60.7%, ahead of Fable 5 (59.9%); Opus 4.8 (legacy) scored 56 |
Not covered in this specific index snapshot |
|
GDPval-AA v2 (agentic knowledge work) |
Elo 1668 |
Opus 5 (max) scores 1861 Elo, more than 100 points ahead of Fable 5 and GPT-5.6 Sol |
Not covered in this specific benchmark |
|
SWE-Bench (harness-dependent) |
67.5% (DeepSWE, KimiCode harness) |
Opus 4.8: 69.2% (Anthropic scaffold) vs 51.9% (Scale AI SEAL); independent Opus 5 figures still emerging |
Comparable roughly 26-point harness spread reported for 3.1 Pro |
|
Terminal-Bench 2.1 |
88.3 (Moonshot-reported), 85% (Artificial Analysis verified) |
84.6 baseline cited for Fable 5/Opus 4.8 in Moonshot's table |
Gemini 3.5 Flash: 76.2% |
|
GPQA Diamond |
Not independently confirmed yet |
Strong; exact Opus 5 figure not yet finalized in public sources |
94.3% (Gemini 3.1 Pro at launch) |
|
Frontend Code Arena (LMArena) |
#1, first in 6 of 7 frontend domains |
Previously held the top spot pre-K3 |
Not leading this board |
Kimi K3
✅ Largest open-weight model released to date; native 1M context; strong agentic/frontend coding results; open weights due July 27, 2026
❌ Newer ecosystem with fewer integrations; less proven long-form writing; higher output-token usage; hallucination rate reportedly rose alongside accuracy gains on at least one benchmark
Claude
✅ Opus 5 now leads independent composite benchmarks at Sonnet-adjacent efficiency; best-in-class writing and instruction-following; most mature developer tooling; predictable flat 1M-context pricing on Sonnet 5; robust enterprise features
❌ Premium tiers (Opus 5, Fable 5) remain the most expensive per-token of the three; smaller context ceiling than Gemini's top claims; no free-tier access to flagship models; Opus 5 still trails Fable 5 and Mythos 5 on the longest-running autonomous and cybersecurity tasks
Gemini
✅ Deepest Google Search and Workspace integration; largest context window across every tier; strong multimodal and scientific reasoning benchmarks; fast, cheap Flash tier
❌ Flagship Gemini 3.5 Pro is delayed and behind internal targets; free tier lost access to Pro-class models as of April 2026; Pro-tier pricing steps up sharply above 200K tokens
|
If you need... |
Best choice |
|
Coding |
Kimi K3 (open-weight, agentic) or Claude Opus 5 (state-of-the-art claims, production reliability) |
|
Writing |
Claude |
|
Research |
Gemini |
|
Students |
Gemini (free tier plus Search) or Claude (structured answers) |
|
Business |
Claude |
|
Google Workspace |
Gemini |
|
Large PDFs |
Gemini (largest context) or Claude (flat pricing on 1M) |
|
Programming |
Kimi K3 or Claude |
|
Creative work |
Claude |
|
Best value |
Kimi K3 |
|
Enterprise |
Claude |
|
Everyday use |
Gemini (free speed) or Claude Sonnet 5 (balanced quality) |
For a deeper dive on Kimi K3, see our What is Moonshot AI Kimi K3: Features, Pricing & Benchmarks.
It depends on the task. Kimi K3 leads on frontend coding and offers a stronger price-to-performance ratio, but Claude's new Opus 5 now leads the Artificial Analysis Intelligence Index overall and remains stronger for writing and verified production coding reliability.
Kimi K3 tends to lead on agentic and coding benchmarks, while Gemini leads on scientific reasoning (GPQA), multimodal understanding (MMMU-Pro), and context window ceiling. Neither is a clean overall winner. It depends on your workload.
Kimi K3 for open-weight, long-horizon agentic coding; Claude Opus 5 for Anthropic's newly claimed state-of-the-art coding results with mature tooling.
Claude, consistently, for tone control, structure, and instruction-following on long-form content.
All three now offer at least 1 million tokens on their flagship tiers, with some Gemini 3.1 Pro sources citing up to 2 million, though that upper figure isn't universally confirmed.
Gemini's Flash tiers are generally fastest for everyday tasks; Kimi K3 tends to use more output tokens per task, which can slow it down in practice.
Gemini and Kimi K3 both offer usable free tiers with quota limits; Claude's free tier gives access to Sonnet 5 but with tighter usage limits than Pro or Max.
Kimi K3 for open-weight deployments; Claude Opus 5 for closed-model users, since it now approaches Fable 5's performance at roughly half the price.
Claude, for tooling maturity (Claude Code, Cowork, IDE integrations) and predictable pricing; Kimi K3 is a strong open-weight alternative for teams that want to self-host or fine-tune.
Claude, thanks to its enterprise plans, Microsoft 365 integration, and Claude Cowork for knowledge-work workflows.
Artificial intelligence is spreading across workplaces faster than many organizations can create the rules needed to control it. The Artificial...
Every time Netflix suggests a show you end up binge-watching, or your email filters out spam without you lifting a...