by Xiaomi
MiMo-V2-Omni is designed for complex real-world multimodal interaction and execution scenarios. We've built an all-modal foundation from the ground up that fuses text, vision, and speech, and use a unified architecture to deeply bind "perception" and "action." This not only breaks the traditional models' limitation of emphasizing understanding over execution, but also natively equips the model with multimodal perception, tool invocation, function execution, and GUI operation capabilities. MiMo-V2-Omni can seamlessly integrate with major agent frameworks, achieving a leap from understanding to manipulation and significantly lowering the barrier to deploying full-modal agents.
MiMo-V2-Omni is designed for complex real-world multimodal interaction and execution scenarios. We've built an all-modal foundation from the ground up that fuses text, vision, and speech, and use a unified architecture to deeply bind "perception" and "action." This not only breaks the traditional models' limitation of emphasizing understanding over execution, but also natively equips the model with multimodal perception, tool invocation, function execution, and GUI operation capabilities. MiMo-V2-Omni can seamlessly integrate with major agent frameworks, achieving a leap from understanding to manipulation and significantly lowering the barrier to deploying full-modal agents.
mimo-v2-omni has a 256,000 token context window.
On AIHubMix, mimo-v2-omni costs $0.44 per million input tokens and $2.2 per million output tokens. Cached input reads are billed at $0.09 per million tokens.
mimo-v2-omni accepts text, image, video and audio input.
mimo-v2-omni supports web. Per-protocol parameter support is listed in the capability table on this page.
mimo-v2-omni is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to mimo-v2-omni — no other code changes needed.
mimo-v2-omni is developed by Xiaomi. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
MiMo-V2.5 is a native, fully multimodal large model designed for agent scenarios; it can…
MiMo-V2.5-Pro is Xiaomi's most powerful model to date. In areas such as general agent…
Only supports OpenAI-compatible formats.
Only supports OpenAI-compatible formats.
Only supports OpenAI-compatible formats.
Only supports OpenAI-compatible formats.
Use mimo-v2-omni via the AIHubMix unified API — one interface for every major LLM.