Qwen/Qwen2.5-VL-72B-Instruct

by Qwen

Qwen2.5-VL is a visual language model from the Qwen2.5 series, equipped with strong visual understanding and reasoning capabilities. It can recognize objects, analyze text and charts, understand key events in long videos, and accurately locate targets within images. The model supports structured output, making it suitable for data such as invoices and forms, and performs excellently in multiple benchmark tests.

API Pricing

Input$0.5 / 1M tokens
Output$0.5 / 1M tokens

Specifications

Modalitiestext, image, video

FAQ

What is Qwen/Qwen2.5-VL-72B-Instruct?

Qwen2.5-VL is a visual language model from the Qwen2.5 series, equipped with strong visual understanding and reasoning capabilities. It can recognize objects, analyze text and charts, understand key events in long videos, and accurately locate targets within images. The model supports structured output, making it suitable for data such as invoices and forms, and performs excellently in multiple benchmark tests.

How much does Qwen/Qwen2.5-VL-72B-Instruct cost?

On AIHubMix, Qwen/Qwen2.5-VL-72B-Instruct costs $0.5 per million input tokens and $0.5 per million output tokens.

What modalities does Qwen/Qwen2.5-VL-72B-Instruct support?

Qwen/Qwen2.5-VL-72B-Instruct accepts text, image and video input.

How do I call Qwen/Qwen2.5-VL-72B-Instruct via API?

Qwen/Qwen2.5-VL-72B-Instruct is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to Qwen/Qwen2.5-VL-72B-Instruct — no other code changes needed.

Who develops Qwen/Qwen2.5-VL-72B-Instruct?

Qwen/Qwen2.5-VL-72B-Instruct is developed by Qwen. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

More from Qwen

qwen3.8-max-preview

by Qwen

Qwen 3.8 Max Preview(Qwen3.8-Max-Preview) is the latest-generation foundation model in…

$0.17/1M in · $0.51/1M out
983,616 tokens context

qwen-audio-3.0-tts-flash

by Qwen

qwen-audio-3.0-tts-flash is a high-performance speech synthesis large model optimized for…

$14.2/1M in · $14.2/1M out

qwen-audio-3.0-tts-plus

by Qwen

qwen-audio-3.0-tts-plus is a high-performance speech synthesis large model designed for…

$15/1M in · $15/1M out

happyhorse-1.1-i2v

by Qwen

HappyHorse-1.1-I2V supports image-to-video generation, further enhancing visual texture…

$2/1M in

happyhorse-1.1-r2v

by Qwen

HappyHorse-1.1-R2V supports reference-based video generation, further improving the…

$2/1M in

happyhorse-1.1-t2v

by Qwen

HappyHorse-1.1-T2V supports text-to-video generation, further enhancing text semantic…

$2/1M in

Use Qwen/Qwen2.5-VL-72B-Instruct via the AIHubMix unified API — one interface for every major LLM.