by Qwen
The model provider is the Sophnet platform. Qwen2-VL-72B-Instruct is the latest iteration in the Qwen2-VL series launched by Alibaba Cloud, representing nearly a year of innovative achievements. This model has 72 billion parameters and can understand images of various resolutions and aspect ratios. Additionally, it supports video understanding of over 20 minutes, enabling high-quality video question answering, dialogue, and content creation, along with complex reasoning and decision-making capabilities. - State-of-the-art image understanding: capable of processing images of various resolutions and aspect ratios, performing excellently across multiple visual understanding benchmarks. - Long video understanding: supports video comprehension exceeding 20 minutes, enabling high-quality video Q&A, dialogues, and content creation. - Agent operation capability: equipped with complex reasoning and decision-making abilities, it can integrate with devices such as phones and robots to perform automated operations based on visual environments and textual instructions. - Multilingual support: in addition to English and Chinese, it supports understanding text in images in multiple languages, including most European languages, Japanese, Korean, Arabic, Vietnamese, and more. - Supports a maximum context length of 128K tokens, offering powerful processing capabilities.
The model provider is the Sophnet platform. Qwen2-VL-72B-Instruct is the latest iteration in the Qwen2-VL series launched by Alibaba Cloud, representing nearly a year of innovative achievements. This model has 72 billion parameters and can understand images of various resolutions and aspect ratios. Additionally, it supports video understanding of over 20 minutes, enabling high-quality video question answering, dialogue, and content creation, along with complex reasoning and decision-making capabilities. - State-of-the-art image understanding: capable of processing images of various resolutions and aspect ratios, performing excellently across multiple visual understanding benchmarks. - Long video understanding: supports video comprehension exceeding 20 minutes, enabling high-quality video Q&A, dialogues, and content creation. - Agent operation capability: equipped with complex reasoning and decision-making abilities, it can integrate with devices such as phones and robots to perform automated operations based on visual environments and textual instructions. - Multilingual support: in addition to English and Chinese, it supports understanding text in images in multiple languages, including most European languages, Japanese, Korean, Arabic, Vietnamese, and more. - Supports a maximum context length of 128K tokens, offering powerful processing capabilities.
On AIHubMix, Qwen2-VL-72B-Instruct costs $2.18 per million input tokens and $6.54 per million output tokens.
Qwen2-VL-72B-Instruct accepts text, image and video input.
Qwen2-VL-72B-Instruct is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to Qwen2-VL-72B-Instruct — no other code changes needed.
Qwen2-VL-72B-Instruct is developed by Qwen. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
Qwen 3.8 Max Preview(Qwen3.8-Max-Preview) is the latest-generation foundation model in…
qwen-audio-3.0-tts-flash is a high-performance speech synthesis large model optimized for…
qwen-audio-3.0-tts-plus is a high-performance speech synthesis large model designed for…
HappyHorse-1.1-I2V supports image-to-video generation, further enhancing visual texture…
HappyHorse-1.1-R2V supports reference-based video generation, further improving the…
HappyHorse-1.1-T2V supports text-to-video generation, further enhancing text semantic…
Use Qwen2-VL-72B-Instruct via the AIHubMix unified API — one interface for every major LLM.