mimo-v2-omni

by Xiaomi

MiMo-V2-Omni is designed for complex real-world multimodal interaction and execution scenarios. We've built an all-modal foundation from the ground up that fuses text, vision, and speech, and use a unified architecture to deeply bind "perception" and "action." This not only breaks the traditional models' limitation of emphasizing understanding over execution, but also natively equips the model with multimodal perception, tool invocation, function execution, and GUI operation capabilities. MiMo-V2-Omni can seamlessly integrate with major agent frameworks, achieving a leap from understanding to manipulation and significantly lowering the barrier to deploying full-modal agents.

API Pricing

Input$0.44 / 1M tokens
Output$2.2 / 1M tokens
Cache read$0.09 / 1M tokens

Specifications

Context window256,000 tokens
Modalitiestext, image, video, audio
Featuresweb

FAQ

What is mimo-v2-omni?

MiMo-V2-Omni is designed for complex real-world multimodal interaction and execution scenarios. We've built an all-modal foundation from the ground up that fuses text, vision, and speech, and use a unified architecture to deeply bind "perception" and "action." This not only breaks the traditional models' limitation of emphasizing understanding over execution, but also natively equips the model with multimodal perception, tool invocation, function execution, and GUI operation capabilities. MiMo-V2-Omni can seamlessly integrate with major agent frameworks, achieving a leap from understanding to manipulation and significantly lowering the barrier to deploying full-modal agents.

What is the context length of mimo-v2-omni?

mimo-v2-omni has a 256,000 token context window.

How much does mimo-v2-omni cost?

On AIHubMix, mimo-v2-omni costs $0.44 per million input tokens and $2.2 per million output tokens. Cached input reads are billed at $0.09 per million tokens.

What modalities does mimo-v2-omni support?

mimo-v2-omni accepts text, image, video and audio input.

What features does mimo-v2-omni support?

mimo-v2-omni supports web. Per-protocol parameter support is listed in the capability table on this page.

How do I call mimo-v2-omni via API?

mimo-v2-omni is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to mimo-v2-omni — no other code changes needed.

Who develops mimo-v2-omni?

mimo-v2-omni is developed by Xiaomi. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

More from Xiaomi

xiaomi-mimo-v2.5

by Xiaomi

MiMo-V2.5 is a native, fully multimodal large model designed for agent scenarios; it can…

$0.15/1M in · $0.31/1M out
256,000 tokens context

xiaomi-mimo-v2.5-pro

by Xiaomi

MiMo-V2.5-Pro is Xiaomi's most powerful model to date. In areas such as general agent…

$0.48/1M in · $0.96/1M out
1,000,000 tokens context

coding-xiaomi-mimo-v2.5

by Xiaomi

Only supports OpenAI-compatible formats.

$0.08/1M in · $0.16/1M out

coding-xiaomi-mimo-v2.5-pro

by Xiaomi

Only supports OpenAI-compatible formats.

$0.2/1M in · $0.4/1M out

coding-xiaomi-mimo-v2-omni

by Xiaomi

Only supports OpenAI-compatible formats.

$0.08/1M in · $0.4/1M out

coding-xiaomi-mimo-v2-pro

by Xiaomi

Only supports OpenAI-compatible formats.

$0.2/1M in · $0.6/1M out

Use mimo-v2-omni via the AIHubMix unified API — one interface for every major LLM.