Qwen2-VL-72B-Instruct

by Qwen

The model provider is the Sophnet platform. Qwen2-VL-72B-Instruct is the latest iteration in the Qwen2-VL series launched by Alibaba Cloud, representing nearly a year of innovative achievements. This model has 72 billion parameters and can understand images of various resolutions and aspect ratios. Additionally, it supports video understanding of over 20 minutes, enabling high-quality video question answering, dialogue, and content creation, along with complex reasoning and decision-making capabilities. - State-of-the-art image understanding: capable of processing images of various resolutions and aspect ratios, performing excellently across multiple visual understanding benchmarks. - Long video understanding: supports video comprehension exceeding 20 minutes, enabling high-quality video Q&A, dialogues, and content creation. - Agent operation capability: equipped with complex reasoning and decision-making abilities, it can integrate with devices such as phones and robots to perform automated operations based on visual environments and textual instructions. - Multilingual support: in addition to English and Chinese, it supports understanding text in images in multiple languages, including most European languages, Japanese, Korean, Arabic, Vietnamese, and more. - Supports a maximum context length of 128K tokens, offering powerful processing capabilities.

API Pricing

Input$2.18 / 1M tokens
Output$6.54 / 1M tokens

Specifications

Modalitiestext, image, video

FAQ

What is Qwen2-VL-72B-Instruct?

The model provider is the Sophnet platform. Qwen2-VL-72B-Instruct is the latest iteration in the Qwen2-VL series launched by Alibaba Cloud, representing nearly a year of innovative achievements. This model has 72 billion parameters and can understand images of various resolutions and aspect ratios. Additionally, it supports video understanding of over 20 minutes, enabling high-quality video question answering, dialogue, and content creation, along with complex reasoning and decision-making capabilities. - State-of-the-art image understanding: capable of processing images of various resolutions and aspect ratios, performing excellently across multiple visual understanding benchmarks. - Long video understanding: supports video comprehension exceeding 20 minutes, enabling high-quality video Q&A, dialogues, and content creation. - Agent operation capability: equipped with complex reasoning and decision-making abilities, it can integrate with devices such as phones and robots to perform automated operations based on visual environments and textual instructions. - Multilingual support: in addition to English and Chinese, it supports understanding text in images in multiple languages, including most European languages, Japanese, Korean, Arabic, Vietnamese, and more. - Supports a maximum context length of 128K tokens, offering powerful processing capabilities.

How much does Qwen2-VL-72B-Instruct cost?

On AIHubMix, Qwen2-VL-72B-Instruct costs $2.18 per million input tokens and $6.54 per million output tokens.

What modalities does Qwen2-VL-72B-Instruct support?

Qwen2-VL-72B-Instruct accepts text, image and video input.

How do I call Qwen2-VL-72B-Instruct via API?

Qwen2-VL-72B-Instruct is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to Qwen2-VL-72B-Instruct — no other code changes needed.

Who develops Qwen2-VL-72B-Instruct?

Qwen2-VL-72B-Instruct is developed by Qwen. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

More from Qwen

qwen3.8-max-preview

by Qwen

Qwen 3.8 Max Preview(Qwen3.8-Max-Preview) is the latest-generation foundation model in…

$0.17/1M in · $0.51/1M out
983,616 tokens context

qwen-audio-3.0-tts-flash

by Qwen

qwen-audio-3.0-tts-flash is a high-performance speech synthesis large model optimized for…

$14.2/1M in · $14.2/1M out

qwen-audio-3.0-tts-plus

by Qwen

qwen-audio-3.0-tts-plus is a high-performance speech synthesis large model designed for…

$15/1M in · $15/1M out

happyhorse-1.1-i2v

by Qwen

HappyHorse-1.1-I2V supports image-to-video generation, further enhancing visual texture…

$2/1M in

happyhorse-1.1-r2v

by Qwen

HappyHorse-1.1-R2V supports reference-based video generation, further improving the…

$2/1M in

happyhorse-1.1-t2v

by Qwen

HappyHorse-1.1-T2V supports text-to-video generation, further enhancing text semantic…

$2/1M in

Use Qwen2-VL-72B-Instruct via the AIHubMix unified API — one interface for every major LLM.