gpt-oss-120b pricing & benchmarks
gpt-oss-120b is available from OpenAI at $0.037 per million input tokens and $0.170 per million output tokens ($0.070 blended at 3:1). It accepts up to 131,072 tokens of context. gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimize…
Pricing
Per million tokens · OpenRouter
Input$0.037
Output$0.170
Blended (3:1)$0.070
Specification
openai/gpt-oss-120b
Context window131K
Max output131K
ReleasedAug 2025
LicenceOpen weights
Modalitiesin:text, out:text
Benchmarks
Independently run — not vendor self-reports. Every score links its source.
| Benchmark | Score | Config | Run | Source |
|---|---|---|---|---|
| GPQA Diamond | 75.8% ±2.7 | high effort | 2025-12-11 | Independently run · Epoch AI |
| SimpleQA Verified | 13.9% ±1.1 | high effort | 2025-12-15 | Independently run · Epoch AI |
| OTIS Mock AIME | 88.9% ±4.5 | high effort | 2025-12-11 | Independently run · Epoch AI |
Scores published by Epoch AI, 'AI Benchmarking Hub', used under CC BY 4.0.
Where to run gpt-oss-120b
20 options across 18 providers · 6.9× between cheapest and dearest. Blended at 3:1. ⚠ marks an option serving less context than the best available, or re-pricing above a prompt-length threshold. Bracketed labels are the provider's own service tier or region.
| Host | Blended $/1M | In / Out | Context | Weights | Uptime 24h |
|---|---|---|---|---|---|
| CoreWeave | $0.065 | $0.030 / $0.170 | 131K | fp4 | 99.2% |
| DeepInfra | $0.070 | $0.037 / $0.170 | 131K | bf16 | 99.9% |
| Novita | $0.100 | $0.050 / $0.250 | 131K | fp4 | 98.4% |
| DigitalOcean | $0.138 | $0.055 / $0.385 | 128K | — | 99.8% |
| SiliconFlow | $0.150 | $0.050 / $0.450 | 131K | fp8 | 87.2% |
| AkashML | $0.150 | $0.037 / $0.490 | 131K | bf16 | 97.3% |
| Google (global) | $0.158 | $0.090 / $0.360 | 131K | — | 91.6% |
| Mancer 2 | $0.193 | $0.090 / $0.500 | 131K | fp8 | 94.1% |
| BaseTen | $0.200 | $0.100 / $0.500 | 128K | fp4 | 99.6% |
| Amazon Bedrock | $0.262 | $0.150 / $0.600 | 131K | — | 99.9% |
| DeepInfra (turbo) | $0.262 | $0.150 / $0.600 | 131K | bf16 | 99.8% |
| Together | $0.262 | $0.150 / $0.600 | 131K | — | 94.4% |
| Nebius | $0.262 | $0.150 / $0.600 | 131K | fp4 | 97.8% |
| Amazon Bedrock (eu-west-1) | $0.262 | $0.150 / $0.600 | 131K | — | — |
| Phala | $0.262 | $0.150 / $0.600 | 131K | — | 99.1% |
| Groq | $0.262 | $0.150 / $0.600 | 131K | — | 100.0% |
| Parasail | $0.263 | $0.100 / $0.750 | 131K | fp4 | 91.5% |
| Mara | $0.300 | $0.150 / $0.750 | 131K | — | 97.2% |
| SambaNova | $0.343 | $0.140 / $0.950 | 131K | — | 94.9% |
| Cerebras | $0.450 | $0.350 / $0.750 | 131K | fp16 | 99.5% |
Host prices and uptime from OpenRouter, synced 2026-08-08. Lower-precision weights (fp4, fp8) can cost less and score worse than the same model at bf16 — the score above is not host-specific.
Compare gpt-oss-120b
FAQ
How much does gpt-oss-120b cost?
gpt-oss-120b costs $0.037 per million input tokens and $0.170 per million output tokens on OpenRouter. At a 3:1 input:output mix that blends to $0.070 per million tokens.
What is gpt-oss-120b's context window?
gpt-oss-120b accepts up to 131,072 tokens of context, and can generate up to 131,072 output tokens.
How good is gpt-oss-120b on benchmarks?
gpt-oss-120b scores 75.8% on GPQA Diamond, independently run and published by Epoch AI. 3 benchmark results are listed on this page, each with its run date and source.