gemini-2.5-flash-lite

by Google

Gemini 2.5 Flash-Lite is a balanced model from Google, optimized for applications that require low-latency performance. It retains the practical capabilities of the Gemini 2.5 family, including configurable reasoning based on budget, integration with tools such as grounding via Google Search and code execution, multimodal input support, and an ultra-long context window of up to 1 million tokens, delivering a strong balance between efficiency, functionality, and cost.

API Pricing

Input$0.1 / 1M tokens
Output$0.4 / 1M tokens
Cache read$0.01 / 1M tokens

Specifications

Context window1,048,576 tokens
Modalitiestext, image, audio, video
Featurestool calling, function calling, structured outputs, long context

FAQ

What is gemini-2.5-flash-lite?

Gemini 2.5 Flash-Lite is a balanced model from Google, optimized for applications that require low-latency performance. It retains the practical capabilities of the Gemini 2.5 family, including configurable reasoning based on budget, integration with tools such as grounding via Google Search and code execution, multimodal input support, and an ultra-long context window of up to 1 million tokens, delivering a strong balance between efficiency, functionality, and cost.

What is the context length of gemini-2.5-flash-lite?

gemini-2.5-flash-lite has a 1,048,576 token context window.

How much does gemini-2.5-flash-lite cost?

On AIHubMix, gemini-2.5-flash-lite costs $0.1 per million input tokens and $0.4 per million output tokens. Cached input reads are billed at $0.01 per million tokens.

What modalities does gemini-2.5-flash-lite support?

gemini-2.5-flash-lite accepts text, image, audio and video input.

What features does gemini-2.5-flash-lite support?

gemini-2.5-flash-lite supports tool calling, function calling, structured outputs and long context. Per-protocol parameter support is listed in the capability table on this page.

How do I call gemini-2.5-flash-lite via API?

gemini-2.5-flash-lite is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to gemini-2.5-flash-lite — no other code changes needed.

Who develops gemini-2.5-flash-lite?

gemini-2.5-flash-lite is developed by Google. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

More from Google

gemini-3.6-flash

by Google

Gemini 3.6 Flash provides sustained frontier-level intelligence optimized for real-world…

$1.5/1M in · $7.5/1M out
1,048,576 tokens context

gemini-3.1-flash-lite-image

by Google

Google's newest, most compact, and most cost-effective image generation and editing…

$0.25/1M in · $1.5/1M out

gemini-3.5-flash-lite

by Google

Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for…

$0.3/1M in · $2.5/1M out
1,048,576 tokens context

gemini-3.5-flash-lite-free

by Google

Gemini 3.5 Flash-Lite free version: Free model resources are limited and provided only…

1,048,576 tokens context

gemini-3.6-flash-free

by Google

Gemini 3.6 Flash free version: fFree model resources are limited and provided only for…

1,000,000 tokens context

gemini-3.5-flash

by Google

Gemini 3.5 Flash provides sustained frontier-level intelligence optimized for real-world…

$1.5/1M in · $9/1M out
1,000,000 tokens context

Use gemini-2.5-flash-lite via the AIHubMix unified API — one interface for every major LLM.