Llama 3.1 8B Instruct pricing & benchmarks
Llama 3.1 8B Instruct is available from Meta at $0.050 per million input tokens and $0.080 per million output tokens ($0.058 blended at 3:1). It accepts up to 131,072 tokens of context. Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...
Pricing
Per million tokens ยท OpenRouter
Input$0.050
Output$0.080
Blended (3:1)$0.058
Specification
meta-llama/llama-3.1-8b-instruct
Context window131K
Max output131K
ReleasedJul 2024
LicenceOpen weights
Modalitiesin:text, out:text
Benchmarks
Independently run โ not vendor self-reports. Every score links its source.
| Benchmark | Score | Config | Run | Source |
|---|---|---|---|---|
| GPQA Diamond | 25.9% ยฑ1.7 | โ | 2025-01-27 | Independently run ยท Epoch AI |
| OTIS Mock AIME | 2.5% ยฑ2.2 | โ | 2025-02-25 | Independently run ยท Epoch AI |
Scores published by Epoch AI, 'AI Benchmarking Hub', used under CC BY 4.0.
Cheaper models that score at least as well on GPQA Diamond
Strictly lower blended price AND an equal-or-higher GPQA Diamond score. One benchmark is one dimension โ a model that wins here may still be weaker at your specific task.
| Model | Provider | Blended $/1M | GPQA Diamond | |
|---|---|---|---|---|
| Mistral Nemo | Mistral AI | $0.022 (โ62%) | 29.9% | Compare โ |
Where to run Llama 3.1 8B Instruct
5 hosts ยท 8.8ร between cheapest and dearest. Blended at 3:1. โ marks an option serving less context than the best available, or re-pricing above a prompt-length threshold. Bracketed labels are the provider's own service tier or region.
| Host | Blended $/1M | In / Out | Context | Weights | Uptime 24h |
|---|---|---|---|---|---|
| DeepInfra | $0.025 | $0.020 / $0.040 | 131K | fp8 | 99.8% |
| Novita โ 16K context, not 131K |
$0.027 | $0.020 / $0.050 | 16K | fp8 | 98.5% |
| Groq | $0.057 | $0.050 / $0.080 | 131K | โ | 99.7% |
| Cloudflare โ 32K context, not 131K |
$0.186 | $0.152 / $0.287 | 32K | fp8 | 100.0% |
| CoreWeave | $0.220 | $0.220 / $0.220 | 128K | bf16 | 100.0% |
Host prices and uptime from OpenRouter, synced 2026-08-08. Lower-precision weights (fp4, fp8) can cost less and score worse than the same model at bf16 โ the score above is not host-specific.
Compare Llama 3.1 8B Instruct
Llama 3.1 8B Instruct vs Qwen3.8 Max
Llama 3.1 8B Instruct vs DeepSeek V4 Flash 0731
Llama 3.1 8B Instruct vs Claude Opus 5
Llama 3.1 8B Instruct vs Gemini 3.6 Flash
Llama 3.1 8B Instruct vs Kimi K3
Llama 3.1 8B Instruct vs GPT-5.6 Luna
Llama 3.1 8B Instruct vs GPT-5.6 Terra
Llama 3.1 8B Instruct vs GPT-5.6 Sol
FAQ
How much does Llama 3.1 8B Instruct cost?
Llama 3.1 8B Instruct costs $0.050 per million input tokens and $0.080 per million output tokens on OpenRouter. At a 3:1 input:output mix that blends to $0.058 per million tokens.
What is Llama 3.1 8B Instruct's context window?
Llama 3.1 8B Instruct accepts up to 131,072 tokens of context, and can generate up to 131,072 output tokens.
How good is Llama 3.1 8B Instruct on benchmarks?
Llama 3.1 8B Instruct scores 25.9% on GPQA Diamond, independently run and published by Epoch AI. 2 benchmark results are listed on this page, each with its run date and source.
Is there a cheaper model as good as Llama 3.1 8B Instruct?
Yes โ 1 model in our catalogue cost less per token than Llama 3.1 8B Instruct and score at least as high on the same benchmark. The cheapest is Mistral Nemo at $0.022 per million tokens blended.