token.appDeepSeek › DeepSeek V3.1 Terminus

DeepSeek V3.1 Terminus pricing & benchmarks

DeepSeek V3.1 Terminus is available from DeepSeek at $0.270 per million input tokens and $1.00 per million output tokens ($0.453 blended at 3:1). It accepts up to 163,840 tokens of context. DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further …

Pricing
Per million tokens · OpenRouter
Input$0.270
Output$1.00
Blended (3:1)$0.453
Specification
deepseek/deepseek-v3.1-terminus
Context window164K
Max output33K
ReleasedSep 2025
LicenceOpen weights
Modalitiesin:text, out:text
No independently-run benchmark scores are published for DeepSeek V3.1 Terminus yet. We show nothing rather than reprinting an unverified figure.
Where to run DeepSeek V3.1 Terminus
5 hosts · 1.2× between cheapest and dearest. Blended at 3:1. ⚠ marks an option serving less context than the best available, or re-pricing above a prompt-length threshold. Bracketed labels are the provider's own service tier or region.
HostBlended $/1MIn / OutContextWeightsUptime 24h
DeepInfra $0.440 $0.270 / $0.950 164K fp4 99.7%
SiliconFlow $0.453 $0.270 / $1.00 164K fp8 97.8%
Novita
⚠ 131K context, not 164K
$0.453 $0.270 / $1.00 131K fp8 100.0%
AtlasCloud
⚠ 131K context, not 164K
$0.462 $0.300 / $0.950 131K fp8 99.7%
StreamLake
⚠ 128K context, not 164K
$0.514 $0.343 / $1.03 128K 99.7%
Host prices and uptime from OpenRouter, synced 2026-08-08. Lower-precision weights (fp4, fp8) can cost less and score worse than the same model at bf16 — the score above is not host-specific.
Compare DeepSeek V3.1 Terminus
FAQ
How much does DeepSeek V3.1 Terminus cost?
DeepSeek V3.1 Terminus costs $0.270 per million input tokens and $1.00 per million output tokens on OpenRouter. At a 3:1 input:output mix that blends to $0.453 per million tokens.
What is DeepSeek V3.1 Terminus's context window?
DeepSeek V3.1 Terminus accepts up to 163,840 tokens of context, and can generate up to 32,768 output tokens.