The Chinese AI developer DeepSeek has released V4 Flash, which it claims is ts most performant and cost-effective model, in a public beta that directly competes with the flagship offerings of Western competitors.
The latest update, V4-Flash-0731, is open source and available via the DeepSeek API. In the coding benchmark Terminal Bench 2.1 it achieved 82.7 per cent, ahead of the 78 per cent scored by Google’s flagship Gemini 3.6 Flash and close to the 85 per cent scored by Anthropic’s Claude Opus 4.8.
DeepSeek has made low API costs a backbone of its product offering and V4 Flash adheres to this rule. The model is accessible for $0.14 per million input tokens and $0.28 per million output tokens.
When instructions are “cached”, or saved in server memory for repeat use, V4 Flash costs $0.0028 per million input tokens. DeepSeek said this makes it cheap to use for repetitive tasks such as coding using a repository or producing documents according to a schema.
In comparison, Gemini 3.6 Flash costs $1.50 per million input tokens and $7.50 per million output tokens and OpenAI’s cheapest frontier model, GPT-5.6 Luna, costs $0.20 per million input tokens and $1.20 per million output tokens.
V4 Flash contains 284 billion parameters, making it lightweight relative to models such as Alibaba’s 2.4 trillion parameter Qwen-3.8-Max and more feasible to run locally with enterprise hardware.
Morgan Linton, co-founder and chief technology officer at the AI data platform Bold Metrics, ran V4 Flash through his custom software engineering benchmark VulcanBench which is designed to test models for on real-world tasks and cost-efficiency.
Run at max performance, the model outclassed Anthropic’s Claude Fable 5 and OpenAI’s GPT-5.6 Sol, passing 91 per cent of the benchmark problems at an average cost of $1.39 compared to $9.78 for Fable 5 and $15.90 for GPT-5.6 Sol.
On low performance settings, V4 Flash still achieved an 89 per cent pass rate, ahead of OpenAI’s best score at a cost of $0.95.
The crowdsourced AI evaluation platform Arena AI also ranks models based on their performance per dollar using the Frontend Code Arena Pareto Frontier, a graph that tracks the balance between an AI model’s performance and API cost.
V4 Flash achieved a score of 1,586 at $0.14 per million input tokens and $0.28 per million output tokens, which Arena AI said makes it “the best performance-per-dollar of any model in its class”.


.jpg)