Home/Models/Llama 4 Scout (109B)
Verified against Meta Llama on Together AI / Groq
Meta / OpenbalancedGenerally Available

Llama 4 Scout (109B) Token Counter & Cost Calculator

Meta open architecture with extreme 10 Million token context window and efficient 17B active MoE

Standard Input / 1M
$0.30
Prompt tokens
Cached Input / 1M
N/A
Not supported
Output / 1M
$0.60
Generation completion
Context Window
10M
Max out: 16.4k

Interactive Cost & Token Simulator for Llama 4 Scout (109B)

Live Calculation
Input Tokens2,500
Output Tokens800
Requests / Day5,000
Cost / Request$0.00123
Daily Spend (5,000 reqs)$6.15
Monthly Run-Rate (30d)$184.50

Standard Workload Cost Scenarios

ScenarioInput TokensOutput TokensUncached CostWith Prompt Caching
Short Chat Query1,000500$0.00060N/A
Document Summarization10,0002,000$0.00420N/A
Codebase & Context Analysis100,00020,000$0.0420N/A
Batch Corpus Processing1,000,000100,000$0.3600N/A
Technical Architecture & Pricing VerificationVerified: September 8, 2026
Model ArchitectureLlama 4 Scout (109B)
Provider OrganizationMeta / Open
Tokenizer EncodingMeta Llama 3/4 Tiktoken (128k vocabulary)
API ConnectivityServed via Together AI, Groq, Fireworks, and self-hosted vLLM/SGLang.
Batch API 50% DiscountNot Available
Tier 1 Rate Quota1,000 RPM · 1M TPM
Blended 3:1 Cost / 1M Tokens$0.38
Open-weights 109B mixture model. Price reflects managed serverless API hosting.Meta Llama on Together AI / Groq

Verified Pricing History

July 2026: $0.45/M input · $0.85/M output ($0.22/M cached)
Llama 4 Scout managed inference pricing baseline.

When to Choose Llama 4 Scout (109B)

Ideal for production workloads demanding balanced capabilities, deep context depth (10M tokens), and reliability from Meta / Open. Excellent when predictable tokenomics and prompt caching support are paramount.

When Another Model May Be Better

If your use-case requires sub-second streaming latency or ultra-high frequency classification at micro-cent pricing, consider lighter budget options such as Gemini Flash-Lite or Claude Haiku. For deep formal logic, consider dedicated reasoning models like o3.

Frequently Asked Questions About Llama 4 Scout (109B)

How much does 1 million tokens cost with Llama 4 Scout (109B)?

For Llama 4 Scout (109B), 1 million input tokens costs $0.30, while 1 million output tokens costs $0.60. If using prompt caching, repetitive input prefixes are discounted to $0.30 per million.

What is the context window for Llama 4 Scout (109B)?

Llama 4 Scout (109B) features a maximum context window of 10,000,000 tokens (~7,500,000 words), with a maximum output limit of 16,384 tokens per completion.

Which tokenizer does Llama 4 Scout (109B) use?

Llama 4 Scout (109B) utilizes the Meta Llama 3/4 Tiktoken (128k vocabulary). Token counting on TokenMath runs client-side to ensure maximum privacy.

Does Llama 4 Scout (109B) support prompt caching discounts?

No, Llama 4 Scout (109B) does not currently advertise prompt caching discounts on its standard API tier.