by InclusionAI
Ling-mini-2.0 is a small-sized, high-performance large language model based on the MoE architecture. It has a total of 16 billion parameters, but only activates 1.4 billion parameters per token (non-embedding 789 million), achieving extremely high generation speed. Thanks to the efficient MoE design and large-scale high-quality training data, despite activating only 1.4 billion parameters, Ling-mini-2.0 still demonstrates top-tier performance on downstream tasks comparable to dense LLMs under 10 billion parameters and even larger-scale MoE models.
Ling-mini-2.0 is a small-sized, high-performance large language model based on the MoE architecture. It has a total of 16 billion parameters, but only activates 1.4 billion parameters per token (non-embedding 789 million), achieving extremely high generation speed. Thanks to the efficient MoE design and large-scale high-quality training data, despite activating only 1.4 billion parameters, Ling-mini-2.0 still demonstrates top-tier performance on downstream tasks comparable to dense LLMs under 10 billion parameters and even larger-scale MoE models.
On AIHubMix, inclusionAI/Ling-mini-2.0 costs $0.07 per million input tokens and $0.27 per million output tokens.
inclusionAI/Ling-mini-2.0 accepts text input.
inclusionAI/Ling-mini-2.0 supports tool calling, function calling and structured outputs. Per-protocol parameter support is listed in the capability table on this page.
inclusionAI/Ling-mini-2.0 is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to inclusionAI/Ling-mini-2.0 — no other code changes needed.
inclusionAI/Ling-mini-2.0 is developed by InclusionAI. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
Ling-1T is the first flagship non-thinking model in the “Ling 2.0” series, featuring 1…
Ring-1T is an open-source idea model with a trillion parameters released by the Bailing…
Ling-flash-2.0 is a language model from inclusionAI with a total of 100 billion…
Ring-flash-2.0 is a high-performance thinking model deeply optimized based on the…
Use inclusionAI/Ling-mini-2.0 via the AIHubMix unified API — one interface for every major LLM.