Anthropic shipped Claude Opus 5 on July 25, 2026, and it is the strongest counter-argument yet to the thesis this site runs on. The numbers first: $5 input / $25 output per million tokens — Opus 4.8's exact price — for a model that now tops the Artificial Analysis Intelligence Index at 61, a point above Claude Fable 5, two above GPT-5.6 Sol, four above Kimi K3. Per AA's task accounting it does this at 26% lower cost than Fable 5. A near-flagship at half a flagship's price is precisely the "capability per dollar" pitch that Chinese models have owned all year. So does the value case survive? Yes — but it moved, and it is worth being precise about where. As of July 26, with day-one caveats flagged.
What shipped
| Claude Opus 5 | |
|---|---|
| Released | July 25, 2026 |
| Price / 1M | $5.00 input · $25.00 output (unchanged from Opus 4.8); Fast mode ≈2.5× speed at 2× price |
| Context | 1M tokens · knowledge cutoff May 2026 |
| Reasoning | Five effort levels — AA scores them 51 (low) to 61 (max), an 8× spread in output tokens |
| AA Intelligence Index | 61 — new #1, ahead of Fable 5 (60), GPT-5.6 Sol (59), Kimi K3 (57) |
| Agentic knowledge work | Tops GDPval-AA v2 and AA-Briefcase; ties #1 on AA's coding-agent index with Claude Code; Terminal-Bench 2.1 ≈ 89%, level with Sol |
Day-one field notes from heavy testing communities, for balance: engineering ability is the consensus win ("the best A-lab model in a year" is a representative verdict); the complaints are verbosity, heavy safety scaffolding, a writing-style regression, and one genuinely counterintuitive finding — several testers rate mid effort as the sweet spot, with max burning tokens for marginal gains. AA's own data adds a sharper caveat: factual-knowledge scores still trail Fable 5, and the hallucination rate on AA-Omniscience jumped 14 points over Opus 4.8, to 50%. The index crown is real; it is not a uniform crown.
The honest math against the Chinese lineup
Run the same reference workload we always use — 60M input + 12M output tokens a month, list prices, no caching:
| Model | Monthly cost | AA Index | Cost vs Opus 5 |
|---|---|---|---|
| Claude Fable 5 | $1,200 | 60 | 2.0× |
| GPT-5.6 Sol | $660 | 59 | 1.1× |
| Claude Opus 5 | $600 | 61 | — |
| Kimi K3 | $360 | 57 | 0.60× |
| GLM-5.2 | $137 | 51 | 0.23× |
| DeepSeek V4-Pro | $37 | — | 0.06× |
Read that table twice, because both readings are true:
What Opus 5 genuinely changes. The closed frontier just repriced itself. Opus 5 makes Fable 5 nearly indefensible for most workloads (same neighborhood of capability, twice the bill) and squeezes GPT-5.6 Sol from above. If your budget lives in the $500+/month band, the calculus between closed options has a new default — and the "Chinese models or bust" version of the value argument is dead. It was already wobbling when K3 priced at $3/$15; Opus 5 confirms the era of lazy 10× gaps is over.
What it does not change. Below that band, the ladder is intact and the gaps are still decisive. Opus 5's output tokens cost 5.7× GLM-5.2's — and output is where agent workloads burn money. GLM-5.2 delivers 84% of Opus 5's index score at 23% of its monthly cost; on coding specifically the gap is 2.5 SWE-bench Pro points against a 4.4× bill. K3 undercuts Opus 5 by 40% on both token prices while sitting four index points back — and K3's own independent cost problem shows Anthropic is not the only lab whose flagship real costs run past sticker. And DeepSeek V4 territory — the $12–37/month tier where most bulk work actually lives — Opus 5 does not even attempt to contest.
The dimension no price cut touches
One more fact from launch week, and it is not a benchmark. This was the week nearly 200 Silicon Valley companies — Y Combinator among them — petitioned the White House not to ban US access to Chinese open models, naming Kimi K3 and Qwen 3.8 — while Nvidia's Jensen Huang, Sundar Pichai and Sam Altman all publicly backed open weights (the full ban question, tracked). The one major lab absent from every one of those letters: Anthropic.
That is the structural difference the price cut cannot touch. Opus 5 is a rental — brilliant, and revocable. GLM-5.2 is MIT-licensed today; K3's weights are committed by July 27; the Chinese lineup can be self-hosted, audited, and kept running whatever policy or pricing does next. For a US startup the open letter is about access risk; for everyone else it is about lock-in. Half-price closed is still closed.
What to actually do
- Frontier-or-nothing workloads (hardest agentic work, deep research): Opus 5 is now the rational default seat — it beats Fable 5's price by half at equal-or-better scores. Genuine credit where due.
- The 90% of work below that bar: the ladder still wins. GLM-5.2 for coding value, K3 (as capacity recovers) for browse-heavy agents, DeepSeek V4 for volume — the full rankings hold.
- Either way: route, don't marry. An OpenAI-compatible endpoint makes "Opus for the hard 10%, Chinese models for the rest" a config file, not an architecture. That blend beats any single-model answer on cost — run yours in the pricing calculator.
Every Chinese model above runs today on Turiloop behind one OpenAI-compatible key — international card, no Chinese phone number.
FAQ
How much does Claude Opus 5 cost? $5 per million input tokens and $25 per million output — the same price as Opus 4.8, for index-topping capability. A Fast mode runs about 2.5× the speed at double the price.
Is Claude Opus 5 better than Kimi K3? On aggregate scoring, yes — AA puts Opus 5 at 61 vs K3's 57, and Opus 5 now leads the agentic knowledge-work boards where K3 was runner-up. K3 counters at 40% lower token prices and open weights committed by July 27. Above ~$500/month budgets Opus 5 is the stronger seat; below it, K3 and GLM-5.2 keep the value crown.
Is Claude Opus 5 better value than GLM-5.2? Per dollar, no. GLM-5.2 posts 84% of Opus 5's index score at roughly 23% of the monthly cost, with output tokens 5.7× cheaper. Opus 5 wins when the task genuinely needs frontier-grade reasoning; GLM-5.2 wins the bill everywhere else.
Does Opus 5 kill the case for Chinese AI models? It kills the lazy version ("closed is always 10× the price"). It does not touch the structural case: open weights, self-hosting, and no exposure to a single vendor's pricing or policy — the week's US open-letter fight named Chinese models as exactly the thing worth protecting access to.