token.app โ€บ Meta โ€บ Llama Guard 4 12B

Llama Guard 4 12B pricing & benchmarks

Llama Guard 4 12B is available from Meta at $0.180 per million input tokens and $0.180 per million output tokens ($0.180 blended at 3:1). It accepts up to 1,048,576 tokens of context. Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM...

Pricing
Per million tokens ยท OpenRouter
Input$0.180
Output$0.180
Blended (3:1)$0.180
Specification
meta-llama/llama-guard-4-12b
Context window1.0M
Max output16K
ReleasedApr 2025
LicenceOpen weights
Modalitiesin:text, in:image, out:text
No independently-run benchmark scores are published for Llama Guard 4 12B yet. We show nothing rather than reprinting an unverified figure.
Where to run Llama Guard 4 12B
2 hosts ยท 1.1ร— between cheapest and dearest. Blended at 3:1. โš  marks an option serving less context than the best available, or re-pricing above a prompt-length threshold. Bracketed labels are the provider's own service tier or region.
HostBlended $/1MIn / OutContextWeightsUptime 24h
DeepInfra
โš  164K context, not 1.0M
$0.180 $0.180 / $0.180 164K bf16 97.4%
Together $0.200 $0.200 / $0.200 1.0M โ€” 99.3%
Host prices and uptime from OpenRouter, synced 2026-08-08. Lower-precision weights (fp4, fp8) can cost less and score worse than the same model at bf16 โ€” the score above is not host-specific.
Compare Llama Guard 4 12B
FAQ
How much does Llama Guard 4 12B cost?
Llama Guard 4 12B costs $0.180 per million input tokens and $0.180 per million output tokens on OpenRouter. At a 3:1 input:output mix that blends to $0.180 per million tokens.
What is Llama Guard 4 12B's context window?
Llama Guard 4 12B accepts up to 1,048,576 tokens of context, and can generate up to 16,384 output tokens.