token.app โ€บ Google โ€บ Gemma 4 31B

Gemma 4 31B pricing & benchmarks

Gemma 4 31B is available from Google at $0.100 per million input tokens and $0.340 per million output tokens ($0.160 blended at 3:1). It accepts up to 262,144 tokens of context. Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Pricing
Per million tokens ยท OpenRouter
Input$0.100
Output$0.340
Blended (3:1)$0.160
Specification
google/gemma-4-31b-it
Context window262K
Max output262K
ReleasedApr 2026
LicenceOpen weights
Modalitiesin:text, in:image, in:video, out:text
Benchmarks
Independently run โ€” not vendor self-reports. Every score links its source.
BenchmarkScoreConfigRunSource
GPQA Diamond 75.8% ยฑ3.1 minimal effort 2026-08-06 Independently run ยท Epoch AI
SimpleQA Verified 9.6% ยฑ0.9 โ€” 2026-04-04 Independently run ยท Epoch AI
OTIS Mock AIME 73.3% ยฑ6.7 minimal effort 2026-08-06 Independently run ยท Epoch AI
Scores published by Epoch AI, 'AI Benchmarking Hub', used under CC BY 4.0.
Cheaper models that score at least as well on GPQA Diamond
Strictly lower blended price AND an equal-or-higher GPQA Diamond score. One benchmark is one dimension โ€” a model that wins here may still be weaker at your specific task.
ModelProviderBlended $/1MGPQA Diamond
gpt-oss-120b OpenAI $0.070 (โˆ’56%) 75.8% Compare โ†’
DeepSeek V4 Flash 0731 DeepSeek $0.113 (โˆ’30%) 91.0% Compare โ†’
Where to run Gemma 4 31B
18 options across 16 providers ยท 7.6ร— between cheapest and dearest. Blended at 3:1. โš  marks an option serving less context than the best available, or re-pricing above a prompt-length threshold. Bracketed labels are the provider's own service tier or region.
HostBlended $/1MIn / OutContextWeightsUptime 24h
OpenInference $0.147 $0.080 / $0.350 262K bf16 92.8%
DeepInfra (turbo) $0.152 $0.090 / $0.340 262K fp4 99.9%
CoreWeave $0.160 $0.100 / $0.340 262K bf16 96.2%
Venice $0.180 $0.120 / $0.360 256K bf16 99.7%
Chutes
โš  131K context, not 262K
$0.182 $0.120 / $0.370 131K fp4 92.1%
DeepInfra $0.193 $0.130 / $0.380 262K fp8 93.3%
SiliconFlow $0.198 $0.130 / $0.400 262K fp8 91.4%
Novita $0.205 $0.140 / $0.400 262K bf16 95.2%
Friendli $0.205 $0.140 / $0.400 262K โ€” 100.0%
Morph
โš  175K context, not 262K
$0.205 $0.140 / $0.400 175K fp4 95.8%
Crusoe $0.205 $0.140 / $0.400 262K โ€” 98.8%
Parasail $0.212 $0.150 / $0.400 262K fp8 96.8%
Phala $0.227 $0.150 / $0.460 262K โ€” 93.5%
ModelRun $0.302 $0.220 / $0.550 262K fp4 96.7%
Together $0.425 $0.280 / $0.860 262K โ€” 85.3%
Together $0.535 $0.390 / $0.970 262K โ€” 95.8%
SambaNova
โš  131K context, not 262K
$0.573 $0.380 / $1.15 131K โ€” 99.7%
Cerebras
โš  131K context, not 262K
$1.11 $0.990 / $1.49 131K fp16 98.8%
Host prices and uptime from OpenRouter, synced 2026-08-08. Lower-precision weights (fp4, fp8) can cost less and score worse than the same model at bf16 โ€” the score above is not host-specific.
Compare Gemma 4 31B
FAQ
How much does Gemma 4 31B cost?
Gemma 4 31B costs $0.100 per million input tokens and $0.340 per million output tokens on OpenRouter. At a 3:1 input:output mix that blends to $0.160 per million tokens.
What is Gemma 4 31B's context window?
Gemma 4 31B accepts up to 262,144 tokens of context, and can generate up to 262,144 output tokens.
How good is Gemma 4 31B on benchmarks?
Gemma 4 31B scores 75.8% on GPQA Diamond, independently run and published by Epoch AI. 3 benchmark results are listed on this page, each with its run date and source.
Is there a cheaper model as good as Gemma 4 31B?
Yes โ€” 2 models in our catalogue cost less per token than Gemma 4 31B and score at least as high on the same benchmark. The cheapest is gpt-oss-120b at $0.070 per million tokens blended.