token.app โ€บ Meta โ€บ Llama 3.3 70B Instruct

Llama 3.3 70B Instruct pricing & benchmarks

Llama 3.3 70B Instruct is available from Meta at $0.100 per million input tokens and $0.320 per million output tokens ($0.155 blended at 3:1). It accepts up to 131,072 tokens of context. The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...

Pricing
Per million tokens ยท OpenRouter
Input$0.100
Output$0.320
Blended (3:1)$0.155
Specification
meta-llama/llama-3.3-70b-instruct
Context window131K
Max output16K
ReleasedDec 2024
LicenceOpen weights
Modalitiesin:text, out:text
Benchmarks
Independently run โ€” not vendor self-reports. Every score links its source.
BenchmarkScoreConfigRunSource
GPQA Diamond 47.4% ยฑ2.9 โ€” 2025-01-27 Independently run ยท Epoch AI
OTIS Mock AIME 5.1% ยฑ2.4 โ€” 2025-02-25 Independently run ยท Epoch AI
Scores published by Epoch AI, 'AI Benchmarking Hub', used under CC BY 4.0.
Cheaper models that score at least as well on GPQA Diamond
Strictly lower blended price AND an equal-or-higher GPQA Diamond score. One benchmark is one dimension โ€” a model that wins here may still be weaker at your specific task.
ModelProviderBlended $/1MGPQA Diamond
gpt-oss-120b OpenAI $0.070 (โˆ’55%) 75.8% Compare โ†’
Phi 4 Microsoft $0.088 (โˆ’44%) 56.1% Compare โ†’
DeepSeek V4 Flash 0731 DeepSeek $0.113 (โˆ’27%) 91.0% Compare โ†’
GPT-5 Nano OpenAI $0.138 (โˆ’11%) 69.4% Compare โ†’
Llama 4 Scout Meta $0.150 (โˆ’3%) 51.8% Compare โ†’
Where to run Llama 3.3 70B Instruct
13 options across 12 providers ยท 6.7ร— between cheapest and dearest. Blended at 3:1. โš  marks an option serving less context than the best available, or re-pricing above a prompt-length threshold. Bracketed labels are the provider's own service tier or region.
HostBlended $/1MIn / OutContextWeightsUptime 24h
DeepInfra (turbo) $0.155 $0.100 / $0.320 131K fp8 94.6%
Nebius $0.198 $0.130 / $0.400 131K fp8 52.6%
AkashML $0.198 $0.130 / $0.400 131K fp8 97.9%
Novita
โš  12K context, not 131K
$0.201 $0.135 / $0.400 12K bf16 97.6%
Parasail $0.290 $0.220 / $0.500 131K fp8 94.3%
Crusoe $0.375 $0.250 / $0.750 131K bf16 99.8%
SambaNova $0.563 $0.450 / $0.900 131K bf16 95.4%
Groq $0.640 $0.590 / $0.790 131K โ€” 99.9%
CoreWeave $0.710 $0.710 / $0.710 128K fp16 99.0%
Google $0.720 $0.720 / $0.720 128K โ€” 95.9%
Google (us-central1) $0.720 $0.720 / $0.720 128K โ€” 100.0%
Cloudflare
โš  24K context, not 131K
$0.783 $0.293 / $2.25 24K fp8 98.0%
Together $1.04 $1.04 / $1.04 131K fp8 94.1%
Host prices and uptime from OpenRouter, synced 2026-08-08. Lower-precision weights (fp4, fp8) can cost less and score worse than the same model at bf16 โ€” the score above is not host-specific.
Compare Llama 3.3 70B Instruct
FAQ
How much does Llama 3.3 70B Instruct cost?
Llama 3.3 70B Instruct costs $0.100 per million input tokens and $0.320 per million output tokens on OpenRouter. At a 3:1 input:output mix that blends to $0.155 per million tokens.
What is Llama 3.3 70B Instruct's context window?
Llama 3.3 70B Instruct accepts up to 131,072 tokens of context, and can generate up to 16,384 output tokens.
How good is Llama 3.3 70B Instruct on benchmarks?
Llama 3.3 70B Instruct scores 47.4% on GPQA Diamond, independently run and published by Epoch AI. 2 benchmark results are listed on this page, each with its run date and source.
Is there a cheaper model as good as Llama 3.3 70B Instruct?
Yes โ€” 5 models in our catalogue cost less per token than Llama 3.3 70B Instruct and score at least as high on the same benchmark. The cheapest is gpt-oss-120b at $0.070 per million tokens blended.