The stable release DeepSeek signaled for mid-July finally started landing on July 31 — half of it, and the half nobody expected to headline. DeepSeek-V4-Flash-0731 entered official public beta with the same 284B-total, 13B-active architecture as the preview, "only re-post-trained" — and the retrain moved it into territory small models are not supposed to reach: 25.2 on Agents' Last Exam, against Claude Opus 4.8's 25.7. A model priced at $0.14/$0.28 per million tokens, within half a point of a $5/$25 flagship on an agentic benchmark. The weights followed to Hugging Face the same day. V4-Pro's stable build is now dated for early August. Here is what shipped, verified, as of August 1.
The numbers
All figures below are DeepSeek's release benchmarks — vendor numbers, pending independent reruns, labeled as such:
| Benchmark | V4-Flash-0731 | Context |
|---|---|---|
| Agents' Last Exam | 25.2 | Opus 4.8 scores 25.7 — near parity at ~1% of the price |
| Terminal-Bench 2.1 | 82.7 | Six points off Kimi K3's 88.3 — from a model 1/10th its size |
| CyberGym | 76.7 | — |
| Toolathlon (verified) | 70.3 | — |
| NL2Repo | 54.2 | — |
| DeepSWE | 54.4 | — |
Coverage framed it bluntly: the retrained Flash beats DeepSeek's own V4-Pro preview on nine agent benchmarks. Community reaction centered on one number — 284B — and one question: how does that fight models five to ten times its size? The answer the release itself suggests: the base model was already good, and the differentiating skill has moved to post-training. Architecture didn't change. Size didn't change. The training recipe did, and it was worth this much.
What actually changed for users
- Agent behavior is the upgrade. The entire benchmark slate is agentic — terminals, tools, repos, cyber ranges. If you run Flash in agent loops, the 0731 build is a different animal from the preview.
- Responses API support, Codex-adapted — the stable build speaks the newer interface out of the box.
- Prices did not move. No change announced: list rates remain $0.14 input / $0.28 output, cache hits near $0.014. The capability moved; the floor price stayed the floor.
- Weights on Hugging Face the same day. Announcement first, upload hours later — but same day. After K3's promise-then-deliver arc and Qwen3.8's still-pending "soon", DeepSeek is the lab that kept release-day openness a fact rather than a roadmap item.
The timing is the message
This landed in the middle of the sharpest price week the model market has had. Twenty-four hours earlier, OpenAI cut GPT-5.6 Luna 80% to $0.20/$1.20 — a shot aimed exactly at the volume tier DeepSeek owns. Flash-0731 is the return fire, and it changes the shape of that fight: Luna competes on price with a model that just posted near-Opus agent scores while staying cheaper on both token rates. The floor of the market is no longer where weak models live; it is where the sharpest engineering shows up.
For the V4 stable-release saga this page has tracked since mid-July, the scorecard reads: slipped past its window, split in two, Flash first — and worth the wait on the evidence. V4-Pro's stable build is dated early August, with a platform notice adjusting Pro's cache pricing from August 3 — the community reads that date as the likely launch day.
What to do with it
- Already routing bulk work to Flash (the pattern our guides have recommended all summer): you just got a free upgrade — same id, same price, better model. On Turiloop,
deepseek-v4-flashrides DeepSeek's official serving line, so the 0731 build lands under the id you already call, no integration change. - Running agent stacks on mid-tier models: re-benchmark. A $0.28-output model at 82.7 Terminal-Bench resets what "good enough for agents" costs — run your own traces before paying flagship rates out of habit. The pricing calculator has the lineup.
- Waiting on V4-Pro: early August, watch August 3. The tracker updates when it lands.
FAQ
What is DeepSeek V4-Flash-0731? The stable release of DeepSeek's V4 Flash, published July 31, 2026 as a public beta: identical 284B/13B MoE architecture to the April preview, fully re-post-trained, with major agent-capability gains (Terminal-Bench 2.1: 82.7; Agents' Last Exam: 25.2) and same-day open weights.
Did DeepSeek V4 Flash prices change? No — list rates remain $0.14 per million input tokens and $0.28 per million output, with cache hits near $0.014. The stable release upgraded capability at unchanged prices.
Is V4-Flash-0731 really close to Claude Opus 4.8? On Agents' Last Exam, DeepSeek's published number is 25.2 vs Opus 4.8's 25.7 — near parity on that one agentic benchmark at roughly 1% of the output price. It is a vendor figure pending independent reruns, and single benchmarks never tell the whole story — but the gap it claims is unprecedented for this size class.
When is DeepSeek V4-Pro's stable release? Early August 2026 per DeepSeek's documentation, with a Pro cache-price adjustment scheduled for August 3 that the community reads as the likely release date.
Do I need to change anything to get the new build? No. The model id is unchanged — deepseek-v4-flash through any OpenAI-compatible endpoint (including Turiloop) serves the stable line as it rolls out.