qwen-audio-3.0-tts-flash

by Qwen

qwen-audio-3.0-tts-flash is a high-performance speech synthesis large model optimized for real-time interactive scenarios. Compared with the previous version, the model supports more low-resource languages and Chinese dialects, improves the authenticity of dialect pronunciation, and enhances free-style instruction following and fine-grained label control, enabling more flexible control of expression such as emotion, tone, character, speaking rate, and volume. At the same time, the model exhibits stronger robustness under complex acoustic conditions like noise and reverberation, improving sound quality, clarity, and overall expressiveness. The Flash version focuses on optimizing the real-time synthesis experience, keeping first-packet latency under 200 ms, making it suitable for low-latency interactive scenarios such as voice assistants, real-time dialogue, and intelligent customer service.

API Pricing

Input$14.2 / 1M tokens
Output$14.2 / 1M tokens

Specifications

Modalitiestext

FAQ

What is qwen-audio-3.0-tts-flash?

qwen-audio-3.0-tts-flash is a high-performance speech synthesis large model optimized for real-time interactive scenarios. Compared with the previous version, the model supports more low-resource languages and Chinese dialects, improves the authenticity of dialect pronunciation, and enhances free-style instruction following and fine-grained label control, enabling more flexible control of expression such as emotion, tone, character, speaking rate, and volume. At the same time, the model exhibits stronger robustness under complex acoustic conditions like noise and reverberation, improving sound quality, clarity, and overall expressiveness. The Flash version focuses on optimizing the real-time synthesis experience, keeping first-packet latency under 200 ms, making it suitable for low-latency interactive scenarios such as voice assistants, real-time dialogue, and intelligent customer service.

How much does qwen-audio-3.0-tts-flash cost?

On AIHubMix, qwen-audio-3.0-tts-flash costs $14.2 per million input tokens and $14.2 per million output tokens.

What modalities does qwen-audio-3.0-tts-flash support?

qwen-audio-3.0-tts-flash accepts text input.

How do I call qwen-audio-3.0-tts-flash via API?

qwen-audio-3.0-tts-flash is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to qwen-audio-3.0-tts-flash — no other code changes needed.

Who develops qwen-audio-3.0-tts-flash?

qwen-audio-3.0-tts-flash is developed by Qwen. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

More from Qwen

qwen3.8-max-preview

by Qwen

Qwen 3.8 Max Preview(Qwen3.8-Max-Preview) is the latest-generation foundation model in…

$0.17/1M in · $0.51/1M out
983,616 tokens context

qwen-audio-3.0-tts-plus

by Qwen

qwen-audio-3.0-tts-plus is a high-performance speech synthesis large model designed for…

$15/1M in · $15/1M out

happyhorse-1.1-i2v

by Qwen

HappyHorse-1.1-I2V supports image-to-video generation, further enhancing visual texture…

$2/1M in

happyhorse-1.1-r2v

by Qwen

HappyHorse-1.1-R2V supports reference-based video generation, further improving the…

$2/1M in

happyhorse-1.1-t2v

by Qwen

HappyHorse-1.1-T2V supports text-to-video generation, further enhancing text semantic…

$2/1M in

qwen3.7-flash

by Qwen

The Qwen 3.7 series' mid-to-high cost-performance "Plus" model builds on strong text…

$0.28/1M in · $1.13/1M out
991,000 tokens context

Use qwen-audio-3.0-tts-flash via the AIHubMix unified API — one interface for every major LLM.