Moonshot said the weights would come by July 27, and on July 27 they came: Kimi K3, all 2.8 trillion parameters, 1.56 terabytes on Hugging Face. In a month when "weights soon" became a launch accessory, Moonshot delivering on the exact committed date deserves to be said plainly — our tracker scored the promise, and the promise held.
Then come the three catches, and they are the useful part of this page: the license is no longer the K2-style Modified MIT; the hardware bill for running it yourself starts around eighty GPUs; and the third-party serving market is thin, slow to commit, and already showing silent quantization. What follows is the honest state of things as of July 30 — what is free, what is gated, and what you should actually do.
The license: free for you, gated for resellers
The community shorthand "license tightened" is directionally right but imprecise. Per the launch coverage of the license text (the Hugging Face license file is the authority):
- Free, no permission needed: downloading K3, running it internally, embedding it in your product, research. If that is you — and it is most companies — the license costs you nothing.
- Gate 1 — MaaS resellers: operate a model-as-a-service business above a $20 million threshold, and you need a separate commercial agreement with Moonshot.
- Gate 2 — attribution at scale: products above 100 million monthly active users or $20 million in monthly revenue must display "Kimi K3" prominently in the interface (the K2 attribution clause, carried over and sharpened).
Read what that targets: not users, not startups — the companies that would resell K3 as an API without paying Moonshot. The community's reading is that this is the Cursor lesson applied; Goldman Sachs framed it the same week as Chinese open models leaving the fully-free era. The strategic point stands either way: open weights are no longer automatically a free lunch for middlemen, and Moonshot has decided that inference margin belongs to Moonshot.
The hardware truth: 1.56TB does not run on wishes
Community deployment reports from the first 72 hours, labeled as such:
- A full-precision-ish deployment demo used 80× RTX 5090 across ten 8-GPU machines, reaching about 20 tokens/second single-stream at MXFP4. The community's verdict on "only 80 GPUs" was appropriately sarcastic.
- Unsloth's 1-bit quantization squeezes K3 to 594GB — it boots on a Mac Studio with 128GB via offloading, but testers put quality retention at roughly 79%. Impressive engineering; not the model that topped the leaderboards.
- Moore Threads announced Day-0 support on domestic MTT S5000 silicon — a 48-card, 3,840GB-VRAM cluster at native MXFP4. Precision claims are unverified, but the geopolitical signal (K3 running end-to-end on non-Nvidia, non-US hardware) is part of why the weights matter.
- Most telling: the big inference hosts that usually race to serve new open models — the Nebius/Cerebras/DeepInfra tier — have mostly hesitated. A 2.8T model with 16-of-896 expert activation is expensive to serve well, and the license now takes a cut of exactly their business model. One early mover listed K3 fast but with output capped at 8K tokens and ~18 tok/s — heavily quantized in practice.
Which yields the consumer-protection rule this month taught: not every endpoint labeled "K3" is the K3 from the benchmarks. Before trusting a third-party K3, ask three questions: what precision is served, what is the real context and output cap, and what throughput do you get at load. Silent quantization is this quarter's quiet scam.
What the release actually changes
Three things, one of them immediately practical:
- The sovereignty hedge is now real, not promised. 1.56TB is on Hugging Face and mirrored beyond it; no policy decision can un-download it. For enterprises that need capability nobody can revoke, K3 joined GLM-5.2 in that category on July 27.
- The validation keeps stacking. The same week, Arena's first full-stack web development leaderboard put K3 (Max) at #1 with 1664, ahead of GPT-5.6 Sol and Claude Fable 5 — a human-preference board, immune to self-reported numbers. Eight of the top twenty are Chinese models; six are open-weight.
- Supply will normalize — through APIs, not home rigs. The official API is already listed on Alibaba's model platform, and downstream gateways stage from there. Self-hosting is for GPU fleets; for everyone else the capacity crunch ends the boring way: more serving capacity, reached through the same OpenAI-compatible interface you already use.
What to actually do
- You have a real GPU fleet and internal workloads: the license-free lane is open. Budget for MXFP4-class deployment and validate quality against the API before trusting any quantization.
- You are anyone else: the API path remains the road. Keep K3 as a model-id behind an OpenAI-compatible endpoint, demand precision disclosures from whoever serves it, and run the pricing math with real output-token counts — K3 thinks expensively.
- Meanwhile: the ladder holds. Kimi K2.6 for long-context agents today, GLM-5.2 for coding value, and K3 slotting in as supply stabilizes — it runs on Turiloop the moment its API supply does, one key, international card.
FAQ
Is Kimi K3 open source now? The weights are public — 1.56TB on Hugging Face since July 27, 2026 — under a custom license: free for internal use, product embedding and research; commercial agreement required for MaaS businesses above $20M; prominent "Kimi K3" attribution required above 100M MAU or $20M monthly revenue.
What hardware do you need to run Kimi K3? Community demos put a usable deployment at roughly 80 datacenter-class GPUs (~20 tok/s at MXFP4). A 1-bit quantization runs on a 128GB Mac Studio at ~79% quality retention — a tech demo, not a production path.
Why are there so few cheap K3 API providers? Serving 2.8T parameters well is expensive, and the new license requires large MaaS resellers to negotiate terms with Moonshot. Both raise the floor price of honest K3 serving — and make silently quantized "K3" endpoints the thing to watch out for.
How do I know a third-party K3 endpoint is really K3? Ask for the served precision, the real context window and output cap, and throughput under load. Cross-check output quality against the official API on your own prompts before committing.
Did Moonshot keep its open-weight promise? Yes — committed by July 27, delivered July 27. The license changed from the K2 era, but the weights are real, complete, and already mirrored.