DeepSeek V4 Flash 0731 pricing & benchmarks
DeepSeek V4 Flash 0731 is available from DeepSeek at $0.090 per million input tokens and $0.180 per million output tokens ($0.113 blended at 3:1). It accepts up to 1,048,576 tokens of context. DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.
Pricing
Per million tokens · OpenRouter
Input$0.090
Output$0.180
Blended (3:1)$0.113
Specification
deepseek/deepseek-v4-flash-0731
Context window1.0M
Max output384K
ReleasedJul 2026
LicenceOpen weights
Modalitiesin:text, out:text
Benchmarks
Independently run — not vendor self-reports. Every score links its source.
| Benchmark | Score | Config | Run | Source |
|---|---|---|---|---|
| GPQA Diamond | 91.0% ±1.8 | max effort | 2026-08-02 | Independently run · Epoch AI |
| FrontierMath T1–3 | 57.5% ±2.9 | max effort | 2026-08-02 | Independently run · Epoch AI |
| FrontierMath T4 | 24.4% ±6.8 | max effort | 2026-08-02 | Independently run · Epoch AI |
| SimpleQA Verified | 34.7% ±1.5 | max effort | 2026-08-02 | Independently run · Epoch AI |
| OTIS Mock AIME | 94.4% ±2.1 | max effort | 2026-08-02 | Independently run · Epoch AI |
Scores published by Epoch AI, 'AI Benchmarking Hub', used under CC BY 4.0.
Where to run DeepSeek V4 Flash 0731
24 hosts · 3.1× between cheapest and dearest. Blended at 3:1. ⚠ marks an option serving less context than the best available, or re-pricing above a prompt-length threshold. Bracketed labels are the provider's own service tier or region.
| Host | Blended $/1M | In / Out | Context | Weights | Uptime 24h |
|---|---|---|---|---|---|
| DeepInfra | $0.113 | $0.090 / $0.180 | 1.0M | fp4 | 99.0% |
| DigitalOcean | $0.158 | $0.126 / $0.252 | 1.0M | — | 98.5% |
| BaseTen | $0.163 | $0.130 / $0.260 | 1.0M | fp8 | 99.8% |
| CoreWeave ⚠ 262K context, not 1.0M |
$0.168 | $0.130 / $0.280 | 262K | fp8 | 99.8% |
| Morph | $0.174 | $0.139 / $0.278 | 1.0M | — | 99.2% |
| Fireworks | $0.175 | $0.140 / $0.280 | 1.0M | — | 92.3% |
| Baidu | $0.175 | $0.140 / $0.280 | 1.0M | fp8 | 99.2% |
| AkashML ⚠ 131K context, not 1.0M |
$0.175 | $0.140 / $0.280 | 131K | fp8 | 98.1% |
| Novita | $0.175 | $0.140 / $0.280 | 1.0M | fp8 | 99.7% |
| Cloudflare ⚠ 384K context, not 1.0M |
$0.175 | $0.140 / $0.280 | 384K | fp8 | 100.0% |
| Together | $0.175 | $0.140 / $0.280 | 1.0M | — | 94.3% |
| DeepSeek | $0.175 | $0.140 / $0.280 | 1.0M | fp8 | 99.8% |
| StreamLake | $0.175 | $0.140 / $0.280 | 1.0M | — | 98.3% |
| Parasail | $0.175 | $0.140 / $0.280 | 1.0M | fp8 | 94.5% |
| AtlasCloud ⚠ 262K context, not 1.0M |
$0.175 | $0.140 / $0.280 | 262K | fp8 | 99.5% |
| GMICloud | $0.175 | $0.140 / $0.280 | 1.0M | fp8 | 100.0% |
| SiliconFlow | $0.175 | $0.140 / $0.280 | 1.0M | fp8 | 98.3% |
| Ionstream | $0.175 | $0.140 / $0.280 | 1.0M | fp4 | 97.6% |
| Ambient | $0.175 | $0.140 / $0.280 | 1.0M | fp4 | 96.6% |
| Io Net ⚠ 262K context, not 1.0M |
$0.200 | $0.160 / $0.320 | 262K | fp8 | 98.8% |
| Venice | $0.219 | $0.175 / $0.350 | 1M | — | 93.7% |
| Phala | $0.250 | $0.200 / $0.400 | 1.0M | — | 97.8% |
| Mancer 2 | $0.256 | $0.175 / $0.500 | 1.0M | fp8 | 84.1% |
| Wafer | $0.350 | $0.280 / $0.560 | 1.0M | — | — |
Host prices and uptime from OpenRouter, synced 2026-08-08. Lower-precision weights (fp4, fp8) can cost less and score worse than the same model at bf16 — the score above is not host-specific.
Compare DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 vs Qwen3.8 Max
DeepSeek V4 Flash 0731 vs Claude Opus 5
DeepSeek V4 Flash 0731 vs Gemini 3.6 Flash
DeepSeek V4 Flash 0731 vs Kimi K3
DeepSeek V4 Flash 0731 vs GPT-5.6 Luna
DeepSeek V4 Flash 0731 vs GPT-5.6 Terra
DeepSeek V4 Flash 0731 vs GPT-5.6 Sol
DeepSeek V4 Flash 0731 vs Grok 4.5
FAQ
How much does DeepSeek V4 Flash 0731 cost?
DeepSeek V4 Flash 0731 costs $0.090 per million input tokens and $0.180 per million output tokens on OpenRouter. At a 3:1 input:output mix that blends to $0.113 per million tokens.
What is DeepSeek V4 Flash 0731's context window?
DeepSeek V4 Flash 0731 accepts up to 1,048,576 tokens of context, and can generate up to 384,000 output tokens.
How good is DeepSeek V4 Flash 0731 on benchmarks?
DeepSeek V4 Flash 0731 scores 91.0% on GPQA Diamond, independently run and published by Epoch AI. 5 benchmark results are listed on this page, each with its run date and source.