PaddleOCR 3.5 brings OCR and document parsing closer to the Hugging Face ecosystem. With this release, supported PaddleOCR models can run with Hugging Face Transformers as an inference backend, so developers working inside PyTorch-based stacks can stop wiring OCR systems together by hand.
What Changed in PaddleOCR 3.5
The release introduces a flexible inference-engine interface. Developers select a backend through the engine parameter and pass backend-specific options through engine_config, meaning the switch happens without changing your existing OCR pipeline.
engineparameter — choosepaddle_static,paddle_dynamic, ortransformersper pipelineengine_config— setdtype,device_type,device_id, andattn_implementationsuch assdpa- Managed pipelines — PaddleOCR still handles every internal component, so you never call modules manually
PaddleOCR continues to ship OCR series like PP-OCRv5 and document parsing series like PaddleOCR-VL 1.5; Transformers joins them as one of several supported runtimes.
Why It Matters for Document AI
For RAG, Document AI, and document-agent applications, the hard part often starts before the LLM. Developers must convert PDFs, scanned pages, tables, charts, formulas, and layouts into reliable structured data — and if that ingestion step is weak, downstream LLM workflows miss context or produce unreliable answers.
With PaddleOCR 3.5, this ingestion layer plugs more naturally into Transformers-centered stacks, reducing integration friction on the path from raw documents to RAG, search, analytics, and automation.
Quick Start
On a CUDA 12.6 environment the setup is a few pip installs and one command:
python -m pip install "paddleocr==3.5.0" "paddlex==3.5.2" "transformers>=5.4.0"
paddleocr ocr -i <image_url> --device gpu:0 --engine transformers
The same interface is available through the Python API, with engine_config used to tune hardware-specific options like running bfloat16 with the scaled-dot-product attention implementation.
When to Use the Transformers Backend
The Transformers backend is a natural fit for teams already relying on Hugging Face infrastructure for model loading, experimentation, and deployment, and for developers who want Hub-compatible model discovery. This release is not about replacing one backend with another — it is about giving developers the freedom to pick the runtime that fits their stack.
What This Means
For the open-source ecosystem, PaddleOCR 3.5 removes a real bottleneck: document ingestion no longer forces a separate runtime just to get text out of images and complex layouts. As more teams build agents and retrieval systems on top of OCR output, having Transformers as a first-class backend lowers the barrier between raw documents and everything downstream of them.