by Qwen
qwen-audio-3.0-tts-plus is a high-performance speech synthesis large model designed for high-quality speech generation scenarios. Compared with the previous version, the model supports more low-resource languages and Chinese dialects, significantly improving the authenticity of dialect pronunciation, and enhancing free-style instruction following and fine-grained label control capabilities, allowing more accurate control of emotion, tone, character, speaking rate, volume, and synthesis style. Meanwhile, the model exhibits stronger robustness under complex acoustic conditions such as noise and reverberation, further improving sound quality, clarity, resolution, and overall expressiveness. The Plus version places greater emphasis on synthesis quality and detail, making it suitable for professional scenarios that require higher sound quality, naturalness, and expressiveness, such as content creation, audiobooks, film and TV dubbing, brand voice design, and high-quality voice services.
qwen-audio-3.0-tts-plus is a high-performance speech synthesis large model designed for high-quality speech generation scenarios. Compared with the previous version, the model supports more low-resource languages and Chinese dialects, significantly improving the authenticity of dialect pronunciation, and enhancing free-style instruction following and fine-grained label control capabilities, allowing more accurate control of emotion, tone, character, speaking rate, volume, and synthesis style. Meanwhile, the model exhibits stronger robustness under complex acoustic conditions such as noise and reverberation, further improving sound quality, clarity, resolution, and overall expressiveness. The Plus version places greater emphasis on synthesis quality and detail, making it suitable for professional scenarios that require higher sound quality, naturalness, and expressiveness, such as content creation, audiobooks, film and TV dubbing, brand voice design, and high-quality voice services.
On AIHubMix, qwen-audio-3.0-tts-plus costs $15 per million input tokens and $15 per million output tokens.
qwen-audio-3.0-tts-plus accepts text input.
qwen-audio-3.0-tts-plus is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to qwen-audio-3.0-tts-plus — no other code changes needed.
qwen-audio-3.0-tts-plus is developed by Qwen. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
Qwen 3.8 Max Preview(Qwen3.8-Max-Preview) is the latest-generation foundation model in…
qwen-audio-3.0-tts-flash is a high-performance speech synthesis large model optimized for…
HappyHorse-1.1-I2V supports image-to-video generation, further enhancing visual texture…
HappyHorse-1.1-R2V supports reference-based video generation, further improving the…
HappyHorse-1.1-T2V supports text-to-video generation, further enhancing text semantic…
The Qwen 3.7 series' mid-to-high cost-performance "Plus" model builds on strong text…
Use qwen-audio-3.0-tts-plus via the AIHubMix unified API — one interface for every major LLM.