Is DeepSeek's V4-Flash Really 100x Cheaper
Than Claude? Inside the Benchmarks

Jul 31 official build · Post-training gains · Harness debut · Kill line · Six-step runbook

DeepSeek V4 Flash 0731 official benchmarks 2026
On July 31, 2026, DeepSeek promoted V4-Flash-0731 to an official API build: same 284B-parameter architecture as April's preview, but a full post-training rerun now beats DeepSeek's own larger V4-Pro preview on agent benchmarks at roughly 1/36 to 1/179 of Claude Opus 4.8 list pricing. If you need a clear timeline, a honest read on whether the scores hold up, and a side-by-side with Kimi K3 and Qwen3.8-Max, this piece covers: ① what actually shipped versus what is still "coming soon"; ② vendor pricing versus Artificial Analysis indices; ③ a six-step adoption runbook through August 5 data.
01

What actually shipped on July 31 — and what didn't

It is easy to read "DeepSeek V4 official" and assume a new model dropped. It did not. The July 31 move is a post-training refresh on the same Flash skeleton. Flagship V4-Pro GA and the in-house Harness agent framework remain unreleased with no firm date.

01

Not a scale-up: DeepSeek states gains came entirely from post-training, not parameter growth — a 284B/13B Flash now beats a 1.6T/49B Pro preview on several agent tasks.

02

API-only rollout: The official build is API public beta only; consumer app and web chat were not updated.

03

Legacy aliases retired: deepseek-chat and deepseek-reasoner were shut down July 24 — unmigrated pipelines fail hard.

04

Competitive pressure: Moonshot's Kimi K3 (2.8T open weights) landed July 27, days before DeepSeek's Flash promotion.

05

Pro GA still vague: Chinese forums mocked "empty promise" delays before Flash surprised; V4-Pro official timing remains unconfirmed.

DateEvent
Apr 24, 2026V4 preview: V4-Pro (1.6T/49B) + V4-Flash (284B/13B), 1M context, MIT weights
Jul 24, 2026Legacy model names retired; traffic routes to V4 family
Jul 27, 2026Kimi K3 2.8T weights on Hugging Face
Jul 31, 2026V4-Flash-0731 official API beta + weights; changelog names Harness "coming soon"
Aug 5, 2026V4-Pro GA still "as soon as possible"; Aug 10–20 windows are unconfirmed rumors
02

DeepSeek V4-Flash pricing vs Kimi K3 and Qwen3.8-Max

ModelStatusTotal / ActiveContextInput miss/hit ($/1M)Output ($/1M)License
V4-Flash-0731Official Jul 31284B / 13B1M$0.14 / $0.0028$0.28MIT
V4-ProPreview1.6T / 49B1M$0.435 / $0.003625$0.87MIT
Kimi K3Weights Jul 272.8T / ~104B est.~1.05M$3.00 / $0.30$15.00Modified MIT
GLM-5.2Open Jun 2026~744B / ~40B1MNot verified hereNot verifiedMIT
Qwen3.8-MaxAPI GA Aug 22.4T / 95B1M$2.00 / ~$0.17–0.25$6.00Weights promised

All figures are vendor-published list prices. DeepSeek has announced a future 2x peak-hour surcharge (9am–12pm and 2pm–6pm Beijing time) with no effective date yet. See our July 20 V4 GA breakdown for the earlier GA narrative; this article focuses on the Jul 31 Flash official build.

Vendor math vs Claude Opus 4.8: about 36x on cache-miss input, 179x on cache-hit input, 89x on output — list prices, not audited totals.

03

How DeepSeek squeezed more from the same model: post-training, CSA+HCA, Harness

V4-Flash-0731 is identical in size and structure to April's preview. DeepSeek attributes the agent benchmark jump entirely to a fresh post-training pass — cutting against the default "bigger equals better" assumption in 2026's model race.

01

Hybrid CSA+HCA attention: Compressed Sparse Attention plus Heavily Compressed Attention, marketed as DSA, cuts compute and memory at long context.

02

mHC hyper-connections: Manifold-Constrained Hyper-Connections refine residual paths for stability.

03

Muon optimizer: Replaces classic optimizers; DeepSeek claims at 1M tokens V4-Pro needs 27% of V3.2 per-token FLOPs and 10% KV cache.

July 31 also marked the first official mention of DeepSeek Harness — an in-house agent runner positioned against Claude Code. Every agent score DeepSeek published for Flash-0731 (Terminal Bench 2.0, Toolathlon, etc.) used Harness minimal mode, which is not yet public. The changelog itself warns agent scores are "extremely sensitive to harness choice."

Note: Terminal Bench 2.0 vendor score: Flash-0731 82.7 vs V4-Pro preview 67.9 — max effort, top_p 0.95, temperature 1.0, Harness minimal mode.

04

China's open-weight melee: Artificial Analysis index and six-step runbook

ModelLabIntelligence IndexCost per task (AA)Note
V4-Flash-0731DeepSeek50$0.03Official Jul 31
Kimi K3Moonshot57$0.86Higher index, ~29x cost
GLM-5.2Zhipu~1 pt above FlashNot verifiedJun 2026 open
GPT-5.6 SolOpenAI9+ pts above$1.86Closed
Claude Fable 5Anthropic9+ pts above$3.15Closed

Intelligence Index and per-task costs cite Artificial Analysis via financial outlets; DeepSeek's own agent scores use different harnesses — do not merge the two columns. DeepSeek is not chasing leaderboard crowns; it is optimizing for "good enough intelligence at a price nobody else matches," which also explains seven weeks atop OpenRouter usage for the preview.

01

Point your client: Set model to deepseek-v4-flash, keep base_url at https://api.deepseek.com; OpenAI and Anthropic SDK shapes both work.

02

Scrub legacy names: Search repos and CI for deepseek-chat and deepseek-reasoner.

03

Staging regression: Switch model names in pre-prod, compare agent pipeline quality, latency, and cache-hit rates.

04

Maximize prompt cache: Pin stable system prompts at the top of messages ($0.0028/M cached input).

05

Route by tier: Flash for bulk and routing; V4-Pro preview for heavy reasoning until official GA.

06

Watch Harness GA: Track changelog and Hugging Face; re-run agent benchmarks once Harness is public.

05

Benchmark caveats, the "kill line," and cite-ready hard data

A

Harness dependency: Agent headline scores are vendor plus unreleased framework — wait for Claude Code / Cursor reproductions.

B

Usability gaps: Overseas dev feedback cites low cache-hit rates and occasional safety-classifier timeouts — compute ceilings still matter.

C

"Kill line" (斩杀线): Chinese dev slang for DeepSeek's good-enough-plus-cheap combo forcing rivals to beat on capability or price — context for GPT-5.6 Luna's 80% cut.

Nvidia, Broadcom, and AMD saw muted moves on July 31 — markets now treat "DeepSeek efficiency" as routine engineering, unlike the early-2025 R1 chip selloff. Reports of a ~$7.4B funding round and IPO prep trace to unnamed financial media sources, not regulatory filings — background only.

Running V4-Flash locally via antirez ds4 needs 96GB unified memory. Pure API usage skips local hardware, but high-frequency agent pipelines still demand stable 24/7 hosts. Consumer Macs and undersized VPS instances swap-thrash under long context; short trials rarely stress 1M-token paths. For steadier iOS CI/CD and agent automation, MESHLAUNCH Mac Mini cloud rental is usually the better production bet: dedicated Apple Silicon, elastic memory tiers, and day/week/month billing.

FAQ

Yes. V4-Pro and V4-Flash, including Jul 31 Flash-0731, ship as MIT weights on Hugging Face for commercial use. Cloud pricing on our rental page.

Vendor list prices: roughly 36x on cache-miss input, 179x on cache-hit input, 89x on output vs Claude Opus 4.8 per million tokens.

No confirmed date. Changelog says "as soon as possible." Aug 10–20 windows are unconfirmed rumors.

SWE-bench Verified and similar third-party suites weigh more. Terminal Bench 2.0 used unreleased Harness — treat as harness-specific until reproduced elsewhere.

DeepSeek's in-house agent runner vs Claude Code — file edits, tool calls, multi-step engineering. Named Jul 31; not yet public. See help center for cloud Mac setup.