by Xiaomi
MiMo-V2-Flash is a mixture of experts (MoE) language model with a total of 309 billion parameters and 15 billion activated parameters. It is designed for high-speed inference and proxy workflows, adopting a novel hybrid attention architecture and multi-token prediction (MTP), significantly reducing inference costs while achieving state-of-the-art performance.
MiMo-V2-Flash is a mixture of experts (MoE) language model with a total of 309 billion parameters and 15 billion activated parameters. It is designed for high-speed inference and proxy workflows, adopting a novel hybrid attention architecture and multi-token prediction (MTP), significantly reducing inference costs while achieving state-of-the-art performance.
On AIHubMix, mimo-v2-flash costs $0.19 per million input tokens and $0.58 per million output tokens. Cached input reads are billed at $0.04 per million tokens.
mimo-v2-flash accepts text input.
mimo-v2-flash supports web. Per-protocol parameter support is listed in the capability table on this page.
mimo-v2-flash is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to mimo-v2-flash — no other code changes needed.
mimo-v2-flash is developed by Xiaomi. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
Free version: mimo-v2-flash-free
MiMo-V2.5 is a native, fully multimodal large model designed for agent scenarios; it can…
MiMo-V2.5-Pro is Xiaomi's most powerful model to date. In areas such as general agent…
Only supports OpenAI-compatible formats.
Only supports OpenAI-compatible formats.
Only supports OpenAI-compatible formats.
Only supports OpenAI-compatible formats.
Use mimo-v2-flash via the AIHubMix unified API — one interface for every major LLM.