Llama 4 Maverick pricing & benchmarks
Llama 4 Maverick is available from Meta at $0.200 per million input tokens and $0.800 per million output tokens ($0.350 blended at 3:1). It accepts up to 1,048,576 tokens of context. Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...
Pricing
Per million tokens ยท OpenRouter
Input$0.200
Output$0.800
Blended (3:1)$0.350
Specification
meta-llama/llama-4-maverick
Context window1.0M
Max output16K
ReleasedApr 2025
LicenceOpen weights
Modalitiesin:text, in:image, out:text
Benchmarks
Independently run โ not vendor self-reports. Every score links its source.
| Benchmark | Score | Config | Run | Source |
|---|---|---|---|---|
| GPQA Diamond | 67.0% ยฑ2.8 | โ | 2025-04-08 | Independently run ยท Epoch AI |
| OTIS Mock AIME | 20.6% ยฑ4.9 | โ | 2025-04-08 | Independently run ยท Epoch AI |
Scores published by Epoch AI, 'AI Benchmarking Hub', used under CC BY 4.0.
Cheaper models that score at least as well on GPQA Diamond
Strictly lower blended price AND an equal-or-higher GPQA Diamond score. One benchmark is one dimension โ a model that wins here may still be weaker at your specific task.
| Model | Provider | Blended $/1M | GPQA Diamond | |
|---|---|---|---|---|
| gpt-oss-120b | OpenAI | $0.070 (โ80%) | 75.8% | Compare โ |
| DeepSeek V4 Flash 0731 | DeepSeek | $0.113 (โ68%) | 91.0% | Compare โ |
| GPT-5 Nano | OpenAI | $0.138 (โ61%) | 69.4% | Compare โ |
| Gemma 4 31B | $0.160 (โ54%) | 75.8% | Compare โ | |
| GPT-5.6 Luna | OpenAI | $0.225 (โ36%) | 91.6% | Compare โ |
Where to run Llama 4 Maverick
5 hosts ยท 1.7ร between cheapest and dearest. Blended at 3:1. โ marks an option serving less context than the best available, or re-pricing above a prompt-length threshold. Bracketed labels are the provider's own service tier or region.
| Host | Blended $/1M | In / Out | Context | Weights | Uptime 24h |
|---|---|---|---|---|---|
| DigitalOcean โ 128K context, not 1.0M |
$0.324 | $0.200 / $0.696 | 128K | โ | 99.9% |
| DeepInfra (base) | $0.350 | $0.200 / $0.800 | 1.0M | fp8 | 99.6% |
| Novita | $0.415 | $0.270 / $0.850 | 1.0M | fp8 | 99.6% |
| Parasail โ 524K context, not 1.0M |
$0.512 | $0.350 / $1.00 | 524K | fp8 | 99.1% |
| Google (us-east5) โ 524K context, not 1.0M |
$0.550 | $0.350 / $1.15 | 524K | โ | 99.9% |
Host prices and uptime from OpenRouter, synced 2026-08-08. Lower-precision weights (fp4, fp8) can cost less and score worse than the same model at bf16 โ the score above is not host-specific.
Compare Llama 4 Maverick
FAQ
How much does Llama 4 Maverick cost?
Llama 4 Maverick costs $0.200 per million input tokens and $0.800 per million output tokens on OpenRouter. At a 3:1 input:output mix that blends to $0.350 per million tokens.
What is Llama 4 Maverick's context window?
Llama 4 Maverick accepts up to 1,048,576 tokens of context, and can generate up to 16,384 output tokens.
How good is Llama 4 Maverick on benchmarks?
Llama 4 Maverick scores 67.0% on GPQA Diamond, independently run and published by Epoch AI. 2 benchmark results are listed on this page, each with its run date and source.
Is there a cheaper model as good as Llama 4 Maverick?
Yes โ 5 models in our catalogue cost less per token than Llama 4 Maverick and score at least as high on the same benchmark. The cheapest is gpt-oss-120b at $0.070 per million tokens blended.