mimo-v2-flash

by Xiaomi

MiMo-V2-Flash is a mixture of experts (MoE) language model with a total of 309 billion parameters and 15 billion activated parameters. It is designed for high-speed inference and proxy workflows, adopting a novel hybrid attention architecture and multi-token prediction (MTP), significantly reducing inference costs while achieving state-of-the-art performance.

API Pricing

Input$0.19 / 1M tokens
Output$0.58 / 1M tokens
Cache read$0.04 / 1M tokens

Specifications

Modalitiestext
Featuresweb

FAQ

What is mimo-v2-flash?

MiMo-V2-Flash is a mixture of experts (MoE) language model with a total of 309 billion parameters and 15 billion activated parameters. It is designed for high-speed inference and proxy workflows, adopting a novel hybrid attention architecture and multi-token prediction (MTP), significantly reducing inference costs while achieving state-of-the-art performance.

How much does mimo-v2-flash cost?

On AIHubMix, mimo-v2-flash costs $0.19 per million input tokens and $0.58 per million output tokens. Cached input reads are billed at $0.04 per million tokens.

What modalities does mimo-v2-flash support?

mimo-v2-flash accepts text input.

What features does mimo-v2-flash support?

mimo-v2-flash supports web. Per-protocol parameter support is listed in the capability table on this page.

How do I call mimo-v2-flash via API?

mimo-v2-flash is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to mimo-v2-flash — no other code changes needed.

Who develops mimo-v2-flash?

mimo-v2-flash is developed by Xiaomi. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

Free version: mimo-v2-flash-free

More from Xiaomi

xiaomi-mimo-v2.5

by Xiaomi

MiMo-V2.5 is a native, fully multimodal large model designed for agent scenarios; it can…

$0.15/1M in · $0.31/1M out
256,000 tokens context

xiaomi-mimo-v2.5-pro

by Xiaomi

MiMo-V2.5-Pro is Xiaomi's most powerful model to date. In areas such as general agent…

$0.48/1M in · $0.96/1M out
1,000,000 tokens context

coding-xiaomi-mimo-v2.5

by Xiaomi

Only supports OpenAI-compatible formats.

$0.08/1M in · $0.16/1M out

coding-xiaomi-mimo-v2.5-pro

by Xiaomi

Only supports OpenAI-compatible formats.

$0.2/1M in · $0.4/1M out

coding-xiaomi-mimo-v2-omni

by Xiaomi

Only supports OpenAI-compatible formats.

$0.08/1M in · $0.4/1M out

coding-xiaomi-mimo-v2-pro

by Xiaomi

Only supports OpenAI-compatible formats.

$0.2/1M in · $0.6/1M out

Use mimo-v2-flash via the AIHubMix unified API — one interface for every major LLM.