If you went looking for a Kimi K3 subscription this week and found nothing to buy — that is not a bug. Moonshot paused new consumer subscriptions roughly 48 hours after K3 launched, citing GPU capacity; the old 199-yuan plan vanished, the new Kimi Code tier promptly showed sold out, and access moved to a reservation queue. And the squeeze is not just on subscriptions: community threads this week document paying API users hitting waves of 429 rate-limit errors, with refund complaints loud enough to trend. Here is what actually happened, what still works, and how to get K3-class capability into your stack today — as of July 25, 2026.
What "sold out" actually means
A quick timeline, all within nine days of launch:
- July 16: K3 ships and immediately tops the charts — #3 on the Artificial Analysis index, second only to Fable 5 on agentic knowledge work.
- July 18: Moonshot pauses new consumer subscriptions, saying demand is pressing against GPU capacity and existing compute goes to current subscribers first.
- The plan shelf gets rebuilt: the 199-yuan tier disappears, replaced by a four-step ladder (49 to 1,399 yuan/month) with Kimi's main products and Kimi Code split — and Kimi Code shows sold out almost immediately.
- July 22: subscriptions begin reopening — but through a reservation queue, not an open buy button.
- July 24: community threads fill with 429 errors on the paid API — one widely-supported complaint thread openly discusses consumer-protection action over paid-but-throttled service.
None of this is a stunt. The demand is real and measurable: independent testing prices K3's agentic work at $10.57 per task, averaging nearly an hour of compute each — a 2.8-trillion-parameter model running a single maximum reasoning effort, at scale, against a fixed GPU pool. Community accounting tells the same story from the user side: one heavy user burned through three Max-tier weekly quotas in days; another spent 540 yuan of API credit on a single deep-research task. Capacity is the product, and K3 consumes it faster than any model Moonshot has shipped.
What this means if you want K3 today
Be precise about what is constrained and what is not:
- Consumer subscriptions: reservation queue. Join it and wait your turn.
- The official API: nominally open, practically rate-limited — the 429 waves are the capacity crunch showing up as errors. Usable for light or off-peak work; painful as a production dependency this week.
- Third-party platforms (Alibaba's Bailian and the API gateways downstream of it): staging K3 gradually rather than opening the floodgates, for the same capacity reasons. This is why K3 is not yet everywhere — including on our own marketplace.
- The weights: committed by July 27 — two days out as this page publishes. If Moonshot delivers, self-hosting and third-party inference providers become the structural escape valve from the capacity bottleneck. That commitment date is now the single most important date on the K3 calendar; the launch tracker is following it.
The practical playbook while K3 recovers
The boring advice is also the correct advice: build against an OpenAI-compatible endpoint now, and treat K3 as a model-id you flip on later. Every serious Chinese flagship speaks the same interface, which turns "waiting for K3" from an architecture problem into a config value.
What covers K3's jobs today, at list prices, without a queue:
- Long-context agent work — Kimi K2.6 ($0.95/$4.00 per million tokens, cached input $0.16): the previous Moonshot flagship, unconstrained and stable, with the same DNA.
- Coding agents — Kimi K2.7-Code, the coding-tuned variant, or GLM-5.2 ($1.40/$4.40, SWE-bench Pro 62.1 — still the value benchmark) for the strongest capability-per-dollar in the lineup.
- Front-end generation — the job K3's testers rave hardest about: GLM-5.2 is the closest available substitute until K3 capacity normalizes.
- Bulk and routine — DeepSeek V4-Flash at $0.14/$0.28; no reason to spend K3-class money here at all.
All of the above run today on Turiloop behind one OpenAI-compatible key — international card, no Chinese phone number — and K3 joins the lineup the moment its API supply stabilizes. Run your own numbers in the pricing calculator. When K3 opens up, switching is one string.
FAQ
Why is Kimi K3 sold out? Demand exceeded Moonshot's GPU capacity within 48 hours of launch. New consumer subscriptions were paused, the plan lineup was rebuilt (49–1,399 yuan tiers), Kimi Code sold out, and access moved to a reservation queue while existing compute serves current subscribers.
Is the Kimi K3 API also affected? Yes, in practice. The API remains open but community reports through July 24 document heavy 429 rate-limiting even for paying users — the same capacity crunch surfacing as errors rather than a closed storefront.
When will Kimi K3 be available again? Moonshot has not published a date for full reopening; subscriptions are trickling back via reservations. The bigger unlock is the open-weight release committed by July 27, 2026, which would let third-party providers serve K3 and relieve the single-vendor bottleneck.
What is the best alternative to Kimi K3 right now? For long-context agents, Kimi K2.6; for coding, GLM-5.2 or Kimi K2.7-Code; for bulk work, DeepSeek V4-Flash. All are OpenAI-compatible, so adopting K3 later is a model-id change, not a migration.
Does "sold out" mean K3 was overhyped? The opposite — the shortage is the strongest demand signal in the July release wave, and independent scoring backs the capability (AA-Briefcase 1543, second only to Fable 5). The constraint is compute economics, not quality.