by Baidu
PaddleOCR-VL is an advanced and efficient document parsing model specifically designed for element recognition within documents. Its core component, PaddleOCR-VL-0.9B, is a compact yet powerful vision-language model (VLM) composed of a NaViT-style dynamic resolution visual encoder and the ERNIE-4.5-0.3B language model, enabling precise element recognition. This model supports 109 languages and excels at recognizing complex elements such as text, tables, formulas, and charts while maintaining extremely low resource consumption. Through comprehensive evaluations on widely used public benchmarks and internal benchmarks, PaddleOCR-VL achieves state-of-the-art (SOTA) performance in both page-level document parsing and element-level recognition. It significantly outperforms existing pipeline-based solutions, multimodal document parsing approaches, and advanced general-purpose multimodal large models, while also offering faster inference speed.
PaddleOCR-VL is an advanced and efficient document parsing model specifically designed for element recognition within documents. Its core component, PaddleOCR-VL-0.9B, is a compact yet powerful vision-language model (VLM) composed of a NaViT-style dynamic resolution visual encoder and the ERNIE-4.5-0.3B language model, enabling precise element recognition. This model supports 109 languages and excels at recognizing complex elements such as text, tables, formulas, and charts while maintaining extremely low resource consumption. Through comprehensive evaluations on widely used public benchmarks and internal benchmarks, PaddleOCR-VL achieves state-of-the-art (SOTA) performance in both page-level document parsing and element-level recognition. It significantly outperforms existing pipeline-based solutions, multimodal document parsing approaches, and advanced general-purpose multimodal large models, while also offering faster inference speed.
On AIHubMix, paddleocr-vl-0.9b costs $2 per million input tokens.
paddleocr-vl-0.9b is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to paddleocr-vl-0.9b — no other code changes needed.
paddleocr-vl-0.9b is developed by Baidu. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
ERNIE 5.1 is the latest model in the Wenxin series, with comprehensive upgrades to its…
ERNIE 5.0 is the next-generation natively multimodal foundation model in the ERNIE…
The Ernie-image-Turbo model is an 8-step distilled version of the Ernie-image model, also…
musesteamer-air-image is a text-to-image model developed by the Baidu Search team aimed…
Qianfan-OCR-Fast is a multimodal large model specialized for OCR, trained primarily on…
Qianfan-OCR-Fast is a multimodal large model specialized for OCR, trained primarily on…
Use paddleocr-vl-0.9b via the AIHubMix unified API — one interface for every major LLM.