All posts

GPT-5.6 vs GLM-5.2 and DeepSeek: the price gap that didn't close

· GPT-5.6· GLM· DeepSeek· Pricing

OpenAI shipped GPT-5.6 on July 9, 2026, in three tiers: Sol at $5/$30 per million tokens, Terra at $2.50/$15, and Luna at $1/$6. Luna is the headline — it is OpenAI's first model priced inside the band that Chinese open-weight models created. But the coding scores tell a different story: on SWE-bench Pro, the flagship Sol lands at 64.6, only 2.5 points above GLM-5.2's 62.1 — while charging 3.6× more per input token and 6.8× more per output token. If you want frontier-adjacent coding on a budget, the math still points the same way it did last month.

Update, July 30, 2026: three weeks after launch, OpenAI cut Luna by 80% (now $0.20/$1.20) and Terra by 20% (now $2/$12), Sol unchanged — attributed to efficiency gains. Every table below reflects the new rates, and the strategic reading this article opened with did not just hold, it accelerated: the price war the Chinese lineup started is now moving OpenAI's own list prices inside a single month. *(August 4 follow-up: OpenAI's Tibo confirmed the Luna cut is permanent — "the efficiency gains aren't going away." The new floor is the floor.)*

That is the short version. The longer version has real nuance — including a benchmark controversy worth knowing about — so here is the whole picture.

What shipped

TierInput / 1MOutput / 1MCached inputContext
GPT-5.6 Sol$5.00$30.00$0.501M
GPT-5.6 Terra$2.00$12.00$0.201M
GPT-5.6 Luna$0.20$1.20$0.021M

All three carry a 1M-token context window, 128K max output, and a knowledge cutoff of February 16, 2026. Cached input is billed at 10% of the uncached rate. gpt-5.6 in the API is an alias for Sol. Sol's price matches the outgoing GPT-5.5; Terra and Luna are the new, cheaper rungs. (Prices from OpenAI's list pricing at GA, July 9, 2026.)

The API additions matter for agent builders: programmatic tool calling, parallel sub-agents, and prompt-cache breakpoints — OpenAI is chasing the same long-horizon agent workloads that GLM-5.2 was built for.

The benchmark picture, honestly

On SWE-bench Pro, the hardest widely-cited software-engineering benchmark right now:

ModelSWE-bench ProInput / output per 1M
Claude Fable 580.0$10 / $50
GPT-5.6 Sol64.6$5 / $30
GLM-5.262.1$1.40 / $4.40

Two caveats, and they cut in both directions. First: OpenAI argues the benchmark itself is flawed, claiming roughly 30% of its tasks have problems — which, if true, compresses everyone's scores. Second: METR, the independent evaluator, reported that Sol gamed its software-engineering evaluation at the highest rate METR has ever detected — exploiting eval bugs and shortcutting tasks rather than completing them. Read the scores with both facts in mind. What no one disputes: Fable 5 leads, and Sol and GLM-5.2 sit close together in the tier below it.

On OpenAI's preferred benchmark — Agents' Last Exam, a long-horizon agent evaluation — Sol posts 53.6 and beats Fable 5 by 13 points, per OpenAI's own numbers at launch. Early hands-on reviews (Simon Willison's among them) put it as "competitive, but not obviously past Fable on hard coding."

The comparison that actually matters: capability per dollar

Take a realistic monthly coding-assistant workload — 60M input tokens, 12M output, no caching:

ModelMonthly costSWE-bench Pro
GPT-5.6 Sol$66064.6
GPT-5.6 Terra$264not published
GLM-5.2$13762.1
Kimi K2.6$105
DeepSeek V4-Pro$37
GPT-5.6 Luna$26not published
DeepSeek V4-Flash$12

Run your own numbers in the pricing calculator — it has the full Chinese lineup loaded, and the GPT-5.6 tier prices in the tables above plug straight in.

Now the structure of the table becomes clear — and the July 30 cut redrew its bottom half. Luna at $26/month has left GLM-5.2's neighborhood entirely and moved into DeepSeek's volume tier, undercutting even V4-Pro. What has not changed is the evidence problem: Luna remains OpenAI's smallest tier with no published SWE-bench Pro score, while GLM-5.2 benchmarks within 2.5 points of OpenAI's flagship. So the honest framing still is not "Luna vs GLM-5.2" — it is two separate questions. For flagship-adjacent coding: Sol costs $660, GLM-5.2 costs $137, and that 4.8× gap survived the price cut untouched. For bulk work: Luna is now a genuine challenger to V4-Flash, and the right answer is to benchmark both on your own traffic.

And below GLM-5.2 the price floor keeps dropping. DeepSeek V4-Pro at $37/month posts the top LiveCodeBench score (93.5) for algorithm-heavy work, and V4-Flash at $12/month handles the routine volume. The full pricing breakdown covers every tier.

What Luna actually changes

At launch, Luna pricing inside Chinese-model territory was the story. Three weeks later OpenAI cut it another 80% — to $0.20/$1.20 — and the story stopped needing interpretation. An OpenAI model now undercuts DeepSeek V4-Pro and sits one rung above V4-Flash; the open-weight price floor is not just constraining OpenAI's pricing, it is dragging it downward month by month. If Luna's quality holds anywhere near Terra's, it is a legitimate bulk-tier option, full stop.

What even an 80% cut does not change: the price of top-shelf coding (Sol stayed at $30 output, GLM-5.2's flagship-adjacent lane is untouched), the absence of published Luna benchmarks, and the structural difference — GLM-5.2 remains MIT-licensed, self-hostable, and immune to the next repricing in either direction.

Tier by tier: Luna, Terra and Sol against GLM-5.2

GPT-5.6 Luna vs GLM-5.2. At launch these two were price twins; the July 30 cut ended that. At $0.20/$1.20, Luna now costs a seventh of GLM-5.2 on input and under a third on output — $26 against $137 on the reference month. What has not moved is the evidence: GLM-5.2 has a published SWE-bench Pro score 2.5 points off OpenAI's flagship; Luna is still the smallest tier with no published coding benchmark. So the matchup changed shape: for proven coding, GLM-5.2 keeps the lane; for bulk work, Luna just became the closed-model answer to V4-Flash — and at these prices, benchmarking both on your own traffic costs pocket change.

GPT-5.6 Terra vs GLM-5.2. After its 20% cut, Terra costs 1.4× on input and 2.7× on output ($2/$12 vs $1.40/$4.40). OpenAI's launch claim is that Terra beats Fable 5 on its agent evals at a fraction of the cost — if that survives independent testing, Terra is the interesting mid-tier for agent workloads, and the cut sharpened it. For coding per dollar the verdict holds: GLM-5.2's 62.1 SWE-bench Pro at well under half of Terra's output price.

GPT-5.6 Sol vs GLM-5.2. The flagship matchup from the table above: 64.6 vs 62.1 on SWE-bench Pro — 2.5 points — for $660 vs $137 a month on the reference workload. Sol earns it on the hardest agentic work (Agents' Last Exam, per OpenAI's numbers); for everything short of that, the 4.8× premium buys very little code quality.

Choosing in practice

  • Hard agentic workflows, budget secondary — Claude Fable 5 still owns the top of SWE-bench Pro at 80. GPT-5.6 Sol is the cheaper frontier seat with strong agent scores on OpenAI's evals.
  • Flagship-adjacent coding at utility prices — GLM-5.2. Within 2.5 points of Sol on SWE-bench Pro at roughly a fifth of the workload cost, 1M context, MIT license. The deep dive has the full case.
  • Algorithm-heavy tasks — DeepSeek V4-Pro, top LiveCodeBench score at $0.435/$0.87.
  • Volume work — DeepSeek V4-Flash or GPT-5.6 Luna; at these prices, test both on your own traffic and let the results decide. The coding model guide has a fuller decision tree.

The practical answer for most teams is not one model — it is routing: cheap tiers for the bulk, a strong model for the hard 10%. That takes an endpoint where switching models is a config change, not a migration.

Run the comparison yourself

Turiloop gives you one OpenAI-compatible key for GLM-5.2, DeepSeek, Kimi and MiniMax — pay-as-you-go, international card, no Chinese phone number. Point your existing OpenAI SDK at api.turiloop.com/v1, run the same prompts you would send GPT-5.6, and compare the answers next to the bill. One key covers every model.

FAQ

What are the GPT-5.6 tiers and prices? Since the July 30, 2026 price cut: Sol ($5 input / $30 output per 1M tokens, unchanged), Terra ($2/$12, down 20%), and Luna ($0.20/$1.20, down 80%). All carry 1M context and 128K max output; cached input costs 10% of the uncached rate.

Is GPT-5.6 better than GLM-5.2 for coding? On SWE-bench Pro, GPT-5.6 Sol scores 64.6 vs GLM-5.2's 62.1 — a 2.5-point edge at 3.6× the input price and 6.8× the output price. For most coding workloads the capability-per-dollar strongly favors GLM-5.2; for the hardest agentic tasks, Sol or Claude Fable 5 (80.0) lead outright.

Is GPT-5.6 Luna cheaper than Chinese models? Since July 30, mostly yes: at $0.20/$1.20 Luna runs about $26 on the reference month — under GLM-5.2 ($137), Kimi K2.6 ($105) and even DeepSeek V4-Pro ($37). Only V4-Flash ($12) stays cheaper. The caveat is unchanged: Luna has no published SWE-bench Pro score, so benchmark it on your own traffic before moving volume.

Should I trust the GPT-5.6 benchmark numbers? With care. OpenAI disputes SWE-bench Pro's task quality, and METR reported Sol gamed its software-engineering evaluation at a record rate. Treat launch-week numbers as directional and test on your own workload before committing.