paddleocr-vl-0.9b

by Baidu

PaddleOCR-VL is an advanced and efficient document parsing model specifically designed for element recognition within documents. Its core component, PaddleOCR-VL-0.9B, is a compact yet powerful vision-language model (VLM) composed of a NaViT-style dynamic resolution visual encoder and the ERNIE-4.5-0.3B language model, enabling precise element recognition. This model supports 109 languages and excels at recognizing complex elements such as text, tables, formulas, and charts while maintaining extremely low resource consumption. Through comprehensive evaluations on widely used public benchmarks and internal benchmarks, PaddleOCR-VL achieves state-of-the-art (SOTA) performance in both page-level document parsing and element-level recognition. It significantly outperforms existing pipeline-based solutions, multimodal document parsing approaches, and advanced general-purpose multimodal large models, while also offering faster inference speed.

API Pricing

Input$2 / 1M tokens

FAQ

What is paddleocr-vl-0.9b?

PaddleOCR-VL is an advanced and efficient document parsing model specifically designed for element recognition within documents. Its core component, PaddleOCR-VL-0.9B, is a compact yet powerful vision-language model (VLM) composed of a NaViT-style dynamic resolution visual encoder and the ERNIE-4.5-0.3B language model, enabling precise element recognition. This model supports 109 languages and excels at recognizing complex elements such as text, tables, formulas, and charts while maintaining extremely low resource consumption. Through comprehensive evaluations on widely used public benchmarks and internal benchmarks, PaddleOCR-VL achieves state-of-the-art (SOTA) performance in both page-level document parsing and element-level recognition. It significantly outperforms existing pipeline-based solutions, multimodal document parsing approaches, and advanced general-purpose multimodal large models, while also offering faster inference speed.

How much does paddleocr-vl-0.9b cost?

On AIHubMix, paddleocr-vl-0.9b costs $2 per million input tokens.

How do I call paddleocr-vl-0.9b via API?

paddleocr-vl-0.9b is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to paddleocr-vl-0.9b — no other code changes needed.

Who develops paddleocr-vl-0.9b?

paddleocr-vl-0.9b is developed by Baidu. AIHubMix aggregates it alongside models from other providers behind one API and one bill.

More from Baidu

ernie-5.1

by Baidu

ERNIE 5.1 is the latest model in the Wenxin series, with comprehensive upgrades to its…

$0.56/1M in · $2.54/1M out
119,000 tokens context

ernie-5.0

by Baidu

ERNIE 5.0 is the next-generation natively multimodal foundation model in the ERNIE…

$0.82/1M in · $3.29/1M out
119,000 tokens context

ernie-image-turbo

by Baidu

The Ernie-image-Turbo model is an 8-step distilled version of the Ernie-image model, also…

$2/1M in

musesteamer-air-image

by Baidu

musesteamer-air-image is a text-to-image model developed by the Baidu Search team aimed…

$2/1M in

qianfan-ocr

by Baidu

Qianfan-OCR-Fast is a multimodal large model specialized for OCR, trained primarily on…

$0.06/1M in · $0.25/1M out
32,000 tokens context

qianfan-ocr-fast

by Baidu

Qianfan-OCR-Fast is a multimodal large model specialized for OCR, trained primarily on…

$0.66/1M in · $2.74/1M out
32,000 tokens context

Use paddleocr-vl-0.9b via the AIHubMix unified API — one interface for every major LLM.