llama-4-maverick

by Llama

Llama 4 Maverick is a high-capacity Mixture-of-Experts (MoE) model from Meta, featuring 400B total parameters and 128 experts, while activating an efficient 17B parameters per inference. Engineered for peak performance, it excels at advanced multimodal tasks. Maverick natively supports text and image input, producing multilingual text and code. With a 1-million-token context window and instruction tuning, it is optimized for complex image reasoning and general-purpose assistant-like interactions. Released under the Llama 4 Community License, Maverick is ideal for research and commercial applications demanding state-of-the-art multimodal understanding and high throughput.

API Pricing

Input$0.2 / 1M tokens
Output$0.2 / 1M tokens

Specifications

Context window1,048,576 tokens
Modalitiestext, image
Featurestool calling, function calling, structured outputs

FAQ

What is llama-4-maverick?

Llama 4 Maverick is a high-capacity Mixture-of-Experts (MoE) model from Meta, featuring 400B total parameters and 128 experts, while activating an efficient 17B parameters per inference. Engineered for peak performance, it excels at advanced multimodal tasks. Maverick natively supports text and image input, producing multilingual text and code. With a 1-million-token context window and instruction tuning, it is optimized for complex image reasoning and general-purpose assistant-like interactions. Released under the Llama 4 Community License, Maverick is ideal for research and commercial applications demanding state-of-the-art multimodal understanding and high throughput.

What is the context length of llama-4-maverick?

llama-4-maverick has a 1,048,576 token context window.

How much does llama-4-maverick cost?

On AIHubMix, llama-4-maverick costs $0.2 per million input tokens and $0.2 per million output tokens.

What modalities does llama-4-maverick support?

llama-4-maverick accepts text and image input.

What features does llama-4-maverick support?

llama-4-maverick supports tool calling, function calling and structured outputs. Per-protocol parameter support is listed in the capability table on this page.

How do I call llama-4-maverick via API?

llama-4-maverick is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to llama-4-maverick — no other code changes needed.

Who develops llama-4-maverick?

llama-4-maverick is developed by Llama. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

More from Llama

llama-4-scout

by Llama

Llama 4 Scout is a highly efficient Mixture-of-Experts (MoE) model from Meta, activating…

$0.2/1M in · $0.2/1M out
131,000 tokens context

llama-3.3-70b

by Llama

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and…

$0.6/1M in · $0.6/1M out
65,536 tokens context

llama-3.1-70b

by Llama
$0.44/1M in · $0.44/1M out

llama3.1-8b

by Llama

cerebras

$0.3/1M in · $0.6/1M out

cerebras-llama-3.3-70b

by Llama
$0.6/1M in · $0.6/1M out

deepinfra-llama-3.1-8b-instant

by Llama
$0.03/1M in · $0.05/1M out

Use llama-4-maverick via the AIHubMix unified API — one interface for every major LLM.