DeepSeek's V4 Flash Exits Preview at Rock-Bottom Prices
DeepSeek's V4 Flash model officially exited preview on July 31 with a sharp jump in coding and agent benchmark scores, while holding its price at just $0.14 per million input tokens — reinforcing how quickly capable AI is getting cheaper.
What changed
V4-Flash-0731 uses the same 284-billion-parameter architecture as the earlier preview version; the update is a retraining, not a new model. But the results moved meaningfully: DeepSeek's own benchmark table shows its Terminal-Bench 2.1 score, a measure of real-world coding and agent task performance, jumping to 82.7 from the preview's 61.8, a 20.9 point gain. It also posted 76.7 on Cybergym and 68.7 on DSBench-FullStack, benchmarks that test security and full-stack development tasks respectively.
Pricing stayed exactly where it was in preview: $0.14 per million input tokens on a cache miss, a fraction of a cent ($0.0028) per million on a cache hit, and $0.28 per million output tokens. That combination, a meaningful capability jump at an unchanged, already-low price, makes V4 Flash one of the cheapest ways to access agent-quality AI performance from any major lab, open-weight or proprietary.
Why this matters to small and medium businesses
If cost has kept you from using AI for coding, data work, or agent-style automation, this is a concrete data point that the price floor keeps dropping. Frontier-adjacent performance at a fraction of a cent per request changes the math on tasks you may have written off as too expensive to automate.
Open-weight models like DeepSeek's can be run through many third-party providers or self-hosted, which gives you more negotiating leverage than relying on a single subscription plan from one of the big-name labs. It's worth knowing this option exists even if you don't switch to it today.
"Cheap" and "unreliable" are no longer the same thing in AI tooling. Benchmark jumps like this one are closing the gap with far more expensive proprietary models fast. If you tested a budget AI tool six months ago and found it lacking, it's worth another look rather than assuming that verdict still holds.
Sources: