token.app › Compare › Llama 3.1 70B Instruct vs Inkling Small

Llama 3.1 70B Instruct vs Inkling Small

Llama 3.1 70B Instruct is 1.6× cheaper per token on a blended 3:1 basis. Inkling Small accepts the larger context window (524K).

Head to head
Blended price uses a 3:1 input:output mix
Llama 3.1 70B InstructInkling Small
ProviderMetaThinkingmachines
Input $/1M$0.400$0.450
Output $/1M$0.400$1.20
Blended $/1M$0.400$0.637
Context window131K524K
Max output16K262K
ReleasedJul 2024Jul 2026
LicenceOpen weightsOpen weights
Choose Llama 3.1 70B Instruct if…
  • Costs less per token — $0.400 vs $0.637 blended
Choose Inkling Small if…
  • Larger context window (524K)
  • Newer model (released Jul 2026)
FAQ
Which is better, Llama 3.1 70B Instruct or Inkling Small?
We do not hold independently-run benchmark scores covering both Llama 3.1 70B Instruct and Inkling Small, so we make no quality claim. Their pricing and specifications are compared on this page.
Is Llama 3.1 70B Instruct cheaper than Inkling Small?
Llama 3.1 70B Instruct is cheaper: $0.400 versus $0.637 per million tokens blended at 3:1 input:output. Input alone: $0.400 vs $0.450. Output alone: $0.400 vs $1.20.
What context windows do Llama 3.1 70B Instruct and Inkling Small have?
Llama 3.1 70B Instruct: 131,072 tokens. Inkling Small: 524,288 tokens.
Who makes Llama 3.1 70B Instruct and Inkling Small?
Llama 3.1 70B Instruct is made by Meta. Inkling Small is made by Thinkingmachines.
Related comparisons