llama-4-scout

by Llama

Llama 4 Scout is a highly efficient Mixture-of-Experts (MoE) model from Meta, activating 17B out of 109B total parameters per inference. It natively supports multimodal input (text and image) and multilingual output (text and code) across 12 languages. Designed for assistant-style interaction and visual reasoning, Scout features a massive 10-million-token context window. It is instruction-tuned for tasks like multilingual chat and image understanding and is released under the Llama 4 Community License for local or commercial deployment.

API Pricing

Input$0.2 / 1M tokens
Output$0.2 / 1M tokens

Specifications

Context window131,000 tokens
Modalitiestext, image
Featurestool calling, function calling, structured outputs

FAQ

What is llama-4-scout?

Llama 4 Scout is a highly efficient Mixture-of-Experts (MoE) model from Meta, activating 17B out of 109B total parameters per inference. It natively supports multimodal input (text and image) and multilingual output (text and code) across 12 languages. Designed for assistant-style interaction and visual reasoning, Scout features a massive 10-million-token context window. It is instruction-tuned for tasks like multilingual chat and image understanding and is released under the Llama 4 Community License for local or commercial deployment.

What is the context length of llama-4-scout?

llama-4-scout has a 131,000 token context window.

How much does llama-4-scout cost?

On AIHubMix, llama-4-scout costs $0.2 per million input tokens and $0.2 per million output tokens.

What modalities does llama-4-scout support?

llama-4-scout accepts text and image input.

What features does llama-4-scout support?

llama-4-scout supports tool calling, function calling and structured outputs. Per-protocol parameter support is listed in the capability table on this page.

How do I call llama-4-scout via API?

llama-4-scout is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to llama-4-scout — no other code changes needed.

Who develops llama-4-scout?

llama-4-scout is developed by Llama. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

More from Llama

llama-4-maverick

by Llama

Llama 4 Maverick is a high-capacity Mixture-of-Experts (MoE) model from Meta, featuring…

$0.2/1M in · $0.2/1M out
1,048,576 tokens context

llama-3.3-70b

by Llama

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and…

$0.6/1M in · $0.6/1M out
65,536 tokens context

llama-3.1-70b

by Llama
$0.44/1M in · $0.44/1M out

llama3.1-8b

by Llama

cerebras

$0.3/1M in · $0.6/1M out

cerebras-llama-3.3-70b

by Llama
$0.6/1M in · $0.6/1M out

deepinfra-llama-3.1-8b-instant

by Llama
$0.03/1M in · $0.05/1M out

Use llama-4-scout via the AIHubMix unified API — one interface for every major LLM.