DeepSeek V4 Flash 0423 pricing & benchmarks
DeepSeek V4 Flash 0423 is available from DeepSeek at $0.140 per million input tokens and $0.280 per million output tokens ($0.175 blended at 3:1). It accepts up to 1,048,576 tokens of context. DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...
Pricing
Per million tokens ยท OpenRouter
Input$0.140
Output$0.280
Blended (3:1)$0.175
Specification
deepseek/deepseek-v4-flash
Context window1.0M
Max output393K
ReleasedApr 2026
LicenceOpen weights
Modalitiesin:text, out:text
No independently-run benchmark scores are published for DeepSeek V4 Flash 0423 yet. We show nothing rather than reprinting an unverified figure.
Where to run DeepSeek V4 Flash 0423
20 hosts ยท 3.2ร between cheapest and dearest. Blended at 3:1. โ marks an option serving less context than the best available, or re-pricing above a prompt-length threshold. Bracketed labels are the provider's own service tier or region.
| Host | Blended $/1M | In / Out | Context | Weights | Uptime 24h |
|---|---|---|---|---|---|
| StreamLake | $0.086 | $0.068 / $0.137 | 1.0M | fp8 | 99.2% |
| Baidu | $0.086 | $0.069 / $0.137 | 1.0M | fp8 | 98.8% |
| OpenInference | $0.098 | $0.070 / $0.180 | 1.0M | fp8 | 95.7% |
| DigitalOcean | $0.105 | $0.084 / $0.168 | 1.0M | โ | 99.3% |
| DeepInfra | $0.113 | $0.090 / $0.180 | 1.0M | fp4 | 99.7% |
| GMICloud | $0.117 | $0.094 / $0.188 | 1.0M | fp8 | 98.5% |
| SiliconFlow | $0.168 | $0.130 / $0.280 | 1.0M | fp8 | 99.1% |
| Alibaba | $0.168 | $0.134 / $0.268 | 1M | fp8 | 98.9% |
| Venice | $0.172 | $0.138 / $0.275 | 1M | โ | 99.0% |
| Morph | $0.174 | $0.139 / $0.278 | 1.0M | โ | 98.2% |
| Parasail | $0.175 | $0.140 / $0.280 | 1.0M | fp8 | 93.7% |
| Fireworks | $0.175 | $0.140 / $0.280 | 1.0M | โ | 98.5% |
| Novita | $0.175 | $0.140 / $0.280 | 1.0M | fp8 | 99.5% |
| Ambient | $0.175 | $0.140 / $0.280 | 1.0M | fp4 | 0.0% |
| Cloudflare โ 384K context, not 1.0M |
$0.175 | $0.140 / $0.280 | 384K | โ | 100.0% |
| AtlasCloud | $0.175 | $0.140 / $0.280 | 1.0M | fp4 | 99.5% |
| CoreWeave | $0.175 | $0.140 / $0.280 | 1.0M | fp8 | 99.3% |
| DeepSeek | $0.175 | $0.140 / $0.280 | 1.0M | โ | 99.8% |
| Phala | $0.250 | $0.200 / $0.400 | 1.0M | โ | 65.4% |
| Mancer 2 | $0.275 | $0.200 / $0.500 | 1.0M | fp4 | 96.3% |
Host prices and uptime from OpenRouter, synced 2026-08-08. Lower-precision weights (fp4, fp8) can cost less and score worse than the same model at bf16 โ the score above is not host-specific.
Subscriptions that include DeepSeek V4 Flash 0423
Consumer plans whose published model list names this model
Compare DeepSeek V4 Flash 0423
DeepSeek V4 Flash 0423 vs Qwen3.8 Max
DeepSeek V4 Flash 0423 vs Claude Opus 5
DeepSeek V4 Flash 0423 vs Gemini 3.6 Flash
DeepSeek V4 Flash 0423 vs Kimi K3
DeepSeek V4 Flash 0423 vs GPT-5.6 Luna
DeepSeek V4 Flash 0423 vs GPT-5.6 Terra
DeepSeek V4 Flash 0423 vs GPT-5.6 Sol
DeepSeek V4 Flash 0423 vs Grok 4.5
FAQ
How much does DeepSeek V4 Flash 0423 cost?
DeepSeek V4 Flash 0423 costs $0.140 per million input tokens and $0.280 per million output tokens on OpenRouter. At a 3:1 input:output mix that blends to $0.175 per million tokens.
What is DeepSeek V4 Flash 0423's context window?
DeepSeek V4 Flash 0423 accepts up to 1,048,576 tokens of context, and can generate up to 393,216 output tokens.