token.app › Compare › Codestral 2508 vs DeepSeek V4 Flash 0731

Codestral 2508 vs DeepSeek V4 Flash 0731

DeepSeek V4 Flash 0731 is 4.0× cheaper per token on a blended 3:1 basis. DeepSeek V4 Flash 0731 accepts the larger context window (1.0M).

Head to head
Blended price uses a 3:1 input:output mix
Codestral 2508DeepSeek V4 Flash 0731
ProviderMistral AIDeepSeek
Input $/1M$0.300$0.090
Output $/1M$0.900$0.180
Blended $/1M$0.450$0.113
Context window256K1.0M
Max output384K
ReleasedAug 2025Jul 2026
LicenceProprietaryOpen weights
Choose Codestral 2508 if…
  • No clear advantage on the data we hold.
Choose DeepSeek V4 Flash 0731 if…
  • Costs less per token — $0.113 vs $0.450 blended
  • Larger context window (1.0M)
  • Open weights — self-hostable
  • Newer model (released Jul 2026)
FAQ
Which is better, Codestral 2508 or DeepSeek V4 Flash 0731?
We do not hold independently-run benchmark scores covering both Codestral 2508 and DeepSeek V4 Flash 0731, so we make no quality claim. Their pricing and specifications are compared on this page.
Is Codestral 2508 cheaper than DeepSeek V4 Flash 0731?
DeepSeek V4 Flash 0731 is cheaper: $0.113 versus $0.450 per million tokens blended at 3:1 input:output. Input alone: $0.300 vs $0.090. Output alone: $0.900 vs $0.180.
What context windows do Codestral 2508 and DeepSeek V4 Flash 0731 have?
Codestral 2508: 256,000 tokens. DeepSeek V4 Flash 0731: 1,048,576 tokens.
Who makes Codestral 2508 and DeepSeek V4 Flash 0731?
Codestral 2508 is made by Mistral AI. DeepSeek V4 Flash 0731 is made by DeepSeek.
Related comparisons