Qwen2-VL-7B-Instruct

by Qwen

The model provider is the Sophnet platform. Qwen2-VL-7B-Instruct is the latest vision-language model launched by Alibaba Cloud and the newest member of the Qwen family. This model is proficient not only in recognizing common objects but also in analyzing text, charts, icons, and layouts within images. As a visual agent, it can reason and dynamically guide tool usage, supporting operations on computers and mobile phones. Additionally, it can understand long videos exceeding one hour and capture key events, accurately locate objects in images, and generate structured outputs for data such as invoices and tables, making it suitable for various scenarios including finance and business. - Vision understanding capability: not only recognizes common objects but also analyzes text, charts, icons, and layouts within images. - Agent capability: functions as a visual agent capable of reasoning and dynamically guiding tool usage, supporting operations on computers and mobile phones. - Long video understanding: can comprehend video content over one hour in length and accurately localize relevant video segments. - Visual localization: precisely locates objects within images by generating bounding boxes or points, providing stable JSON coordinate outputs. - Structured output: supports structured data output for invoices, tables, and other data, suitable for finance, business, and various other scenarios.

API Pricing

Input$0.28 / 1M tokens
Output$0.7 / 1M tokens

Specifications

Modalitiestext, image, video

FAQ

What is Qwen2-VL-7B-Instruct?

The model provider is the Sophnet platform. Qwen2-VL-7B-Instruct is the latest vision-language model launched by Alibaba Cloud and the newest member of the Qwen family. This model is proficient not only in recognizing common objects but also in analyzing text, charts, icons, and layouts within images. As a visual agent, it can reason and dynamically guide tool usage, supporting operations on computers and mobile phones. Additionally, it can understand long videos exceeding one hour and capture key events, accurately locate objects in images, and generate structured outputs for data such as invoices and tables, making it suitable for various scenarios including finance and business. - Vision understanding capability: not only recognizes common objects but also analyzes text, charts, icons, and layouts within images. - Agent capability: functions as a visual agent capable of reasoning and dynamically guiding tool usage, supporting operations on computers and mobile phones. - Long video understanding: can comprehend video content over one hour in length and accurately localize relevant video segments. - Visual localization: precisely locates objects within images by generating bounding boxes or points, providing stable JSON coordinate outputs. - Structured output: supports structured data output for invoices, tables, and other data, suitable for finance, business, and various other scenarios.

How much does Qwen2-VL-7B-Instruct cost?

On AIHubMix, Qwen2-VL-7B-Instruct costs $0.28 per million input tokens and $0.7 per million output tokens.

What modalities does Qwen2-VL-7B-Instruct support?

Qwen2-VL-7B-Instruct accepts text, image and video input.

How do I call Qwen2-VL-7B-Instruct via API?

Qwen2-VL-7B-Instruct is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to Qwen2-VL-7B-Instruct — no other code changes needed.

Who develops Qwen2-VL-7B-Instruct?

Qwen2-VL-7B-Instruct is developed by Qwen. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

More from Qwen

qwen3.8-max-preview

by Qwen

Qwen 3.8 Max Preview(Qwen3.8-Max-Preview) is the latest-generation foundation model in…

$0.17/1M in · $0.51/1M out
983,616 tokens context

qwen-audio-3.0-tts-flash

by Qwen

qwen-audio-3.0-tts-flash is a high-performance speech synthesis large model optimized for…

$14.2/1M in · $14.2/1M out

qwen-audio-3.0-tts-plus

by Qwen

qwen-audio-3.0-tts-plus is a high-performance speech synthesis large model designed for…

$15/1M in · $15/1M out

happyhorse-1.1-i2v

by Qwen

HappyHorse-1.1-I2V supports image-to-video generation, further enhancing visual texture…

$2/1M in

happyhorse-1.1-r2v

by Qwen

HappyHorse-1.1-R2V supports reference-based video generation, further improving the…

$2/1M in

happyhorse-1.1-t2v

by Qwen

HappyHorse-1.1-T2V supports text-to-video generation, further enhancing text semantic…

$2/1M in

Use Qwen2-VL-7B-Instruct via the AIHubMix unified API — one interface for every major LLM.