qwen3-vl-flash

by Qwen

The Qwen3 series of compact visual-understanding models achieves an effective fusion of thinking mode and non-thinking mode, outperforming the open-source Qwen3-VL-30B-A3B with faster response speeds. It comprehensively upgrades image and video understanding, supporting ultra-long contexts such as long videos and long documents, spatial awareness, and universal object recognition; it also possesses visual 2D/3D localization capabilities and is capable of handling complex real-world tasks.

API Pricing

Input$0.02 / 1M tokens
Output$0.21 / 1M tokens
Cache read$0.0041 / 1M tokens

Specifications

Context window254,000 tokens
Modalitiestext, image, video
Featurestool calling, function calling, structured outputs

FAQ

What is qwen3-vl-flash?

The Qwen3 series of compact visual-understanding models achieves an effective fusion of thinking mode and non-thinking mode, outperforming the open-source Qwen3-VL-30B-A3B with faster response speeds. It comprehensively upgrades image and video understanding, supporting ultra-long contexts such as long videos and long documents, spatial awareness, and universal object recognition; it also possesses visual 2D/3D localization capabilities and is capable of handling complex real-world tasks.

What is the context length of qwen3-vl-flash?

qwen3-vl-flash has a 254,000 token context window.

How much does qwen3-vl-flash cost?

On AIHubMix, qwen3-vl-flash costs $0.02 per million input tokens and $0.21 per million output tokens. Cached input reads are billed at $0.0041 per million tokens.

What modalities does qwen3-vl-flash support?

qwen3-vl-flash accepts text, image and video input.

What features does qwen3-vl-flash support?

qwen3-vl-flash supports tool calling, function calling and structured outputs. Per-protocol parameter support is listed in the capability table on this page.

How do I call qwen3-vl-flash via API?

qwen3-vl-flash is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to qwen3-vl-flash — no other code changes needed.

Who develops qwen3-vl-flash?

qwen3-vl-flash is developed by Qwen. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

More from Qwen

qwen3.8-max-preview

by Qwen

Qwen 3.8 Max Preview(Qwen3.8-Max-Preview) is the latest-generation foundation model in…

$0.17/1M in · $0.51/1M out
983,616 tokens context

qwen-audio-3.0-tts-flash

by Qwen

qwen-audio-3.0-tts-flash is a high-performance speech synthesis large model optimized for…

$14.2/1M in · $14.2/1M out

qwen-audio-3.0-tts-plus

by Qwen

qwen-audio-3.0-tts-plus is a high-performance speech synthesis large model designed for…

$15/1M in · $15/1M out

happyhorse-1.1-i2v

by Qwen

HappyHorse-1.1-I2V supports image-to-video generation, further enhancing visual texture…

$2/1M in

happyhorse-1.1-r2v

by Qwen

HappyHorse-1.1-R2V supports reference-based video generation, further improving the…

$2/1M in

happyhorse-1.1-t2v

by Qwen

HappyHorse-1.1-T2V supports text-to-video generation, further enhancing text semantic…

$2/1M in

Use qwen3-vl-flash via the AIHubMix unified API — one interface for every major LLM.