All posts

DeepSeek V4 Flash goes stable: 284B parameters, punching at Opus weight

· DeepSeek· V4 Flash· Release

The stable release DeepSeek signaled for mid-July finally started landing on July 31 — half of it, and the half nobody expected to headline. DeepSeek-V4-Flash-0731 entered official public beta with the same 284B-total, 13B-active architecture as the preview, "only re-post-trained" — and the retrain moved it into territory small models are not supposed to reach: 25.2 on Agents' Last Exam, against Claude Opus 4.8's 25.7. A model priced at $0.14/$0.28 per million tokens, within half a point of a $5/$25 flagship on an agentic benchmark. The weights followed to Hugging Face the same day. V4-Pro's stable build is now dated for early August. Here is what shipped, verified, as of August 1.

The numbers

All figures below are DeepSeek's release benchmarks — vendor numbers, pending independent reruns, labeled as such:

BenchmarkV4-Flash-0731Context
Agents' Last Exam25.2Opus 4.8 scores 25.7 — near parity at ~1% of the price
Terminal-Bench 2.182.7Six points off Kimi K3's 88.3 — from a model 1/10th its size
CyberGym76.7
Toolathlon (verified)70.3
NL2Repo54.2
DeepSWE54.4

Coverage framed it bluntly: the retrained Flash beats DeepSeek's own V4-Pro preview on nine agent benchmarks. Community reaction centered on one number — 284B — and one question: how does that fight models five to ten times its size? The answer the release itself suggests: the base model was already good, and the differentiating skill has moved to post-training. Architecture didn't change. Size didn't change. The training recipe did, and it was worth this much.

What actually changed for users

  • Agent behavior is the upgrade. The entire benchmark slate is agentic — terminals, tools, repos, cyber ranges. If you run Flash in agent loops, the 0731 build is a different animal from the preview.
  • Responses API support, Codex-adapted — the stable build speaks the newer interface out of the box.
  • Prices did not move. No change announced: list rates remain $0.14 input / $0.28 output, cache hits near $0.014. The capability moved; the floor price stayed the floor.
  • Weights on Hugging Face the same day. Announcement first, upload hours later — but same day. After K3's promise-then-deliver arc and Qwen3.8's still-pending "soon", DeepSeek is the lab that kept release-day openness a fact rather than a roadmap item.

The timing is the message

This landed in the middle of the sharpest price week the model market has had. Twenty-four hours earlier, OpenAI cut GPT-5.6 Luna 80% to $0.20/$1.20 — a shot aimed exactly at the volume tier DeepSeek owns. Flash-0731 is the return fire, and it changes the shape of that fight: Luna competes on price with a model that just posted near-Opus agent scores while staying cheaper on both token rates. The floor of the market is no longer where weak models live; it is where the sharpest engineering shows up.

For the V4 stable-release saga this page has tracked since mid-July, the scorecard reads: slipped past its window, split in two, Flash first — and worth the wait on the evidence. V4-Pro's stable build is dated early August, with a platform notice adjusting Pro's cache pricing from August 3 — the community reads that date as the likely launch day.

What to do with it

  • Already routing bulk work to Flash (the pattern our guides have recommended all summer): you just got a free upgrade — same id, same price, better model. On Turiloop, deepseek-v4-flash rides DeepSeek's official serving line, so the 0731 build lands under the id you already call, no integration change.
  • Running agent stacks on mid-tier models: re-benchmark. A $0.28-output model at 82.7 Terminal-Bench resets what "good enough for agents" costs — run your own traces before paying flagship rates out of habit. The pricing calculator has the lineup.
  • Waiting on V4-Pro: early August, watch August 3. The tracker updates when it lands.

FAQ

What is DeepSeek V4-Flash-0731? The stable release of DeepSeek's V4 Flash, published July 31, 2026 as a public beta: identical 284B/13B MoE architecture to the April preview, fully re-post-trained, with major agent-capability gains (Terminal-Bench 2.1: 82.7; Agents' Last Exam: 25.2) and same-day open weights.

Did DeepSeek V4 Flash prices change? No — list rates remain $0.14 per million input tokens and $0.28 per million output, with cache hits near $0.014. The stable release upgraded capability at unchanged prices.

Is V4-Flash-0731 really close to Claude Opus 4.8? On Agents' Last Exam, DeepSeek's published number is 25.2 vs Opus 4.8's 25.7 — near parity on that one agentic benchmark at roughly 1% of the output price. It is a vendor figure pending independent reruns, and single benchmarks never tell the whole story — but the gap it claims is unprecedented for this size class.

When is DeepSeek V4-Pro's stable release? Early August 2026 per DeepSeek's documentation, with a Pro cache-price adjustment scheduled for August 3 that the community reads as the likely release date.

Do I need to change anything to get the new build? No. The model id is unchanged — deepseek-v4-flash through any OpenAI-compatible endpoint (including Turiloop) serves the stable line as it rolls out.