by InclusionAI
Ling-flash-2.0 is a language model from inclusionAI with a total of 100 billion parameters, of which 6.1 billion are activated per token (4.8 billion non-embedding). As part of the Ling 2.0 architecture series, it is designed as a lightweight yet powerful Mixture-of-Experts (MoE) model. It aims to deliver performance comparable to or even exceeding that of 40B-level dense models and other larger MoE models, but with a significantly smaller active parameter count. The model represents a strategy focused on achieving high performance and efficiency through extreme architectural design and training methods.
Ling-flash-2.0 is a language model from inclusionAI with a total of 100 billion parameters, of which 6.1 billion are activated per token (4.8 billion non-embedding). As part of the Ling 2.0 architecture series, it is designed as a lightweight yet powerful Mixture-of-Experts (MoE) model. It aims to deliver performance comparable to or even exceeding that of 40B-level dense models and other larger MoE models, but with a significantly smaller active parameter count. The model represents a strategy focused on achieving high performance and efficiency through extreme architectural design and training methods.
On AIHubMix, inclusionAI/Ling-flash-2.0 costs $0.14 per million input tokens and $0.54 per million output tokens.
inclusionAI/Ling-flash-2.0 accepts text input.
inclusionAI/Ling-flash-2.0 supports tool calling, function calling and structured outputs. Per-protocol parameter support is listed in the capability table on this page.
inclusionAI/Ling-flash-2.0 is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to inclusionAI/Ling-flash-2.0 — no other code changes needed.
inclusionAI/Ling-flash-2.0 is developed by InclusionAI. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
Ling-1T is the first flagship non-thinking model in the “Ling 2.0” series, featuring 1…
Ring-1T is an open-source idea model with a trillion parameters released by the Bailing…
Ling-mini-2.0 is a small-sized, high-performance large language model based on the MoE…
Ring-flash-2.0 is a high-performance thinking model deeply optimized based on the…
Use inclusionAI/Ling-flash-2.0 via the AIHubMix unified API — one interface for every major LLM.