All posts

Qwen3.8-Max is official: 2.4T parameters, $2/$6 — and the first Max-class open weights

· Qwen· Release· Alibaba

Alibaba made Qwen3.8-Max official on August 3, 2026, and answered almost every question this site has been tracking since the July 19 preview. The spec sheet is now real: 2.4 trillion total parameters, roughly 95B active per token, a 1M-token context window (991K max input, 131K max output), native text-plus-vision input, and API pricing of $2 per million input tokens, $6 per million output, $0.25 cached. And then the sentence nobody was betting on: the weights go open next week — both the 2.4T flagship and a distilled Qwen3.8-27B — on Hugging Face and ModelScope. No Max-class Qwen has ever been open-sourced. The lab whose previous flagship promised "soon" and stayed closed just committed to shipping the largest open model on the market. Here is what shipped, verified, as of August 4.

The numbers, and who they aim at

All scores below are Alibaba's release benchmarks — vendor numbers, pending independent reruns. Alibaba's own materials note that external models were evaluated on their preferred harnesses and some benchmarks are Qwen's own; treat cross-lab comparisons as directional.

BenchmarkQwen3.8-MaxContext
Terminal-Bench 2.186.6Above Claude Opus 4.8's 84.6 — and DeepSeek V4-Flash's 82.7
SWE-bench Pro67.7Opus 4.8 scores 69.2; Claude Fable 5 scores 80.0
PaperBench93.0Above GPT-5.6 Sol's 90.5
FrontierSWE73.5Above Opus 4.8's 70.0
GPQA Diamond92.6
OSWorld-Verified86.1
DeepSWE 1.156.6

One set of numbers is not vendor-controlled: the Arenas. Qwen3.8-Max entered WebDev Arena at #4 (1668 Elo) — Opus 5 Max holds #1 at 1705 — Text Arena at #5 (1496), and Vision Arena at #2 (1305). Crowd-judged rankings have their own biases, but they are nobody's marketing department, and top-5 across three boards is not something a weak model does.

The shape of the pitch is clear from the slate: terminals, repos, long-horizon agent work, research tasks. Alibaba's demos leaned the same way — multi-day autonomous coding runs, 500-turn chip-design sessions. This is a model aimed at the workload class where Kimi K3 planted its flag in July, at half of K3's output price.

The open-weights turn

Two weeks ago, our preview tracker filed Qwen3.8's open-source promise under "pattern risk": Qwen3.6-Max made similar noises and never opened, and we wrote that announcing openness buys the halo while delivering it is a separate decision. That skepticism now has an expiry date. The commitment is specific — next week, weights on Hugging Face and ModelScope, both the 2.4T Max and a 27B distillation — and it comes with an obvious motive the community named instantly: Kimi K3's open weights were taking the "largest open model" crown uncontested. One widely-echoed read on the Chinese forums: *K3 forced Alibaba's hand — the open-source throne was slipping.*

Two practical notes before the halo settles:

  • A 2.4T open model is open in license, not in practice. Serving it means multi-GPU nodes in the 8×H100/B300 class at minimum. For nearly everyone, the artifact that matters is Qwen3.8-27B — that is the one that will actually run on hardware people own. The pattern of K3's weights release applies unchanged: the flagship weights are for auditability, negotiating leverage and hosting providers; the small model is for you.
  • The license is unnamed until the upload exists. July's lesson, twice over: K3 delivered weights on the promised day but swapped Modified MIT for a MaaS-gated license on the way. What terms Qwen3.8 ships under — and whether "next week" holds — is the remaining open item, and this page updates when it lands.

The honest two readings

The bull case. Terminal-Bench above Opus 4.8, an Arena top-5 sweep, native multimodality (vision folded into the agent loop, at an estimated fraction of a cent per image), a 1M context, and $2/$6 pricing that undercuts every Western flagship tier and Kimi K3 alike. If even most of the vendor slate survives independent reruns, this is the most complete frontier package any Chinese lab has shipped.

The bear case. The preview's two-week record taught caution, and the launch didn't erase it. The community verdict that stuck — *"never loses a benchmark, never wins a real test"* — was earned across two weeks of mixed hands-on results. The famous pelican-drawing test finally produced a stunning image — after 12 minutes of thinking. Early testers consistently flag the same trade: scores bought with enormous reasoning budgets, which bills as output tokens and reads as latency. And at $2/$6, it costs roughly 40% more than GLM-5.2 ($1.40/$4.40), which remains within a hair of the best coding scores per dollar. Qwen3.8-Max has to justify the premium on multimodality and long-horizon stamina — which is exactly the ground where vendor demos are least verifiable.

Both readings fit one sentence: the specs are frontier-class and public now; the efficiency is the part nobody has independently measured.

Where the price lands

ModelInput / Output per M tokensNotes
DeepSeek V4-Flash$0.14 / $0.28The floor, freshly retrained
GLM-5.2$1.40 / $4.40Value default, MIT weights
Qwen3.8-Max$2 / $61M context, native vision, cache $0.25
GPT-5.6 Terra$2 / $12Post-July-30 cut
Kimi K3$3 / $15Priciest Chinese lab model
Claude Opus 5$5 / $25Western flagship price floor

The slot is deliberate: above the value tier, below every flagship — with a context window and multimodal loop the cheaper Chinese options don't have. August's price war just gained a middleweight.

What to do with it

  • The API is live now at $2/$6 — no Token Plan bundles required, unlike the preview weeks. On Turiloop, Qwen3.7-Max and Qwen3.7-Plus serve today behind one OpenAI-compatible key (international card, no Chinese phone number), and Qwen3.8 joins the lineup as its API listing rolls out — switching will be a model-id change. The pricing calculator picks it up the same day.
  • If your workload is agentic and long-horizon, put it on the shortlist against K3 and V4-Flash-0731 — and benchmark with your own traces, because every model in this class right now scores better on vendor slates than in someone else's harness.
  • If you're waiting on the weights: next week, allegedly. The tracker logs delivery — or the slip — either way.

FAQ

What is Qwen3.8-Max? Alibaba's flagship model, officially released August 3, 2026 after a two-week preview: a 2.4-trillion-parameter MoE with ~95B active parameters per token, 1M-token context, native text+vision input, priced at $2/$6 per million tokens with $0.25 cached input.

Is Qwen3.8 open source? Not yet — but the commitment is now specific: weights for both Qwen3.8-Max (2.4T) and Qwen3.8-27B are promised on Hugging Face and ModelScope within a week of the August 3 release. It would be the first Max-class Qwen ever opened. License terms are still unnamed.

How does Qwen3.8-Max compare to Kimi K3? On vendor benchmarks they trade blows — K3 holds Terminal-Bench 2.1 (88.3 vs 86.6) while Qwen3.8 counters on research and OS-control tasks — but Qwen3.8-Max costs $2/$6 against K3's $3/$15, less than half the output price. K3's numbers have had a month of independent scrutiny; Qwen3.8's are days old.

Is Qwen3.8-Max good at coding? Vendor numbers say strong but not top: SWE-bench Pro 67.7 sits below Opus 4.8 (69.2) and well below Fable 5 (80.0), while Terminal-Bench 2.1 (86.6) is genuinely elite. Community testing during the preview found real quality with heavy thinking-time cost. For pure coding value, GLM-5.2 at $1.40/$4.40 remains the default recommendation.

How do I access Qwen3.8-Max outside China? The official API is live at $2/$6. Through OpenAI-compatible gateways like Turiloop, the current Qwen lineup (3.7-Max, 3.7-Plus) works today with an international card and no Chinese phone number — see the Qwen access guide — and 3.8 lands as its listing opens.