Llama Guard 4 12B pricing & benchmarks
Llama Guard 4 12B is available from Meta at $0.180 per million input tokens and $0.180 per million output tokens ($0.180 blended at 3:1). It accepts up to 1,048,576 tokens of context. Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM...
Pricing
Per million tokens ยท OpenRouter
Input$0.180
Output$0.180
Blended (3:1)$0.180
Specification
meta-llama/llama-guard-4-12b
Context window1.0M
Max output16K
ReleasedApr 2025
LicenceOpen weights
Modalitiesin:text, in:image, out:text
No independently-run benchmark scores are published for Llama Guard 4 12B yet. We show nothing rather than reprinting an unverified figure.
Where to run Llama Guard 4 12B
2 hosts ยท 1.1ร between cheapest and dearest. Blended at 3:1. โ marks an option serving less context than the best available, or re-pricing above a prompt-length threshold. Bracketed labels are the provider's own service tier or region.
| Host | Blended $/1M | In / Out | Context | Weights | Uptime 24h |
|---|---|---|---|---|---|
| DeepInfra โ 164K context, not 1.0M |
$0.180 | $0.180 / $0.180 | 164K | bf16 | 97.4% |
| Together | $0.200 | $0.200 / $0.200 | 1.0M | โ | 99.3% |
Host prices and uptime from OpenRouter, synced 2026-08-08. Lower-precision weights (fp4, fp8) can cost less and score worse than the same model at bf16 โ the score above is not host-specific.
Compare Llama Guard 4 12B
FAQ
How much does Llama Guard 4 12B cost?
Llama Guard 4 12B costs $0.180 per million input tokens and $0.180 per million output tokens on OpenRouter. At a 3:1 input:output mix that blends to $0.180 per million tokens.
What is Llama Guard 4 12B's context window?
Llama Guard 4 12B accepts up to 1,048,576 tokens of context, and can generate up to 16,384 output tokens.