by Google
This latest 2.5 Flash model comes with improvements in two key areas we heard consistent feedback on: Better agentic tool use: We've improved how the model uses tools, leading to better performance in more complex, agentic and multi-step applications. This model shows noticeable improvements on key agentic benchmarks, including a 5% gain on SWE-Bench Verified, compared to our last release (48.9% → 54%). More efficient: With thinking on, the model is now significantly more cost-efficient—achieving higher quality outputs while using fewer tokens, reducing latency and cost (see charts above).
This latest 2.5 Flash model comes with improvements in two key areas we heard consistent feedback on: Better agentic tool use: We've improved how the model uses tools, leading to better performance in more complex, agentic and multi-step applications. This model shows noticeable improvements on key agentic benchmarks, including a 5% gain on SWE-Bench Verified, compared to our last release (48.9% → 54%). More efficient: With thinking on, the model is now significantly more cost-efficient—achieving higher quality outputs while using fewer tokens, reducing latency and cost (see charts above).
gemini-2.5-flash-preview-09-2025 has a 1,048,576 token context window.
On AIHubMix, gemini-2.5-flash-preview-09-2025 costs $0.3 per million input tokens and $2.5 per million output tokens. Cached input reads are billed at $0.03 per million tokens.
gemini-2.5-flash-preview-09-2025 accepts text, image, audio and video input.
gemini-2.5-flash-preview-09-2025 supports tool calling, function calling and structured outputs. Per-protocol parameter support is listed in the capability table on this page.
gemini-2.5-flash-preview-09-2025 is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to gemini-2.5-flash-preview-09-2025 — no other code changes needed.
gemini-2.5-flash-preview-09-2025 is developed by Google. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
Gemini 3.6 Flash provides sustained frontier-level intelligence optimized for real-world…
Google's newest, most compact, and most cost-effective image generation and editing…
Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for…
Gemini 3.5 Flash-Lite free version: Free model resources are limited and provided only…
Gemini 3.6 Flash free version: fFree model resources are limited and provided only for…
Gemini 3.5 Flash provides sustained frontier-level intelligence optimized for real-world…
Use gemini-2.5-flash-preview-09-2025 via the AIHubMix unified API — one interface for every major LLM.