PaddleOCR 3.5: Running OCR and Document Parsing Tasks with a Transformers Backend

PaddleOCR 3.5 adds Hugging Face Transformers as a document parsing inference backend, letting PP-OCRv5 and PaddleOCR-VL 1.5 models run wherever the engine parameter is set.

Monday May 18, 2026 Source: huggingface.co
TL;DR — Quick Answer

PaddleOCR 3.5, shared May 18 2026, lets supported OCR and document parsing models such as PP-OCRv5 and PaddleOCR-VL 1.5 run with Hugging Face Transformers as an inference backend. Developers switch backends through the engine parameter and tune options such as dtype, device placement, and attention implementation through engine_config. PaddlePaddle still manages the OCR pipeline behind the scenes, while Transformers fits the stack into PyTorch-based RAG, Document AI, search, and agent workflows.

Key Takeaways

PaddleOCR 3.5: Running OCR and Document Parsing Tasks with a Transformers Backend — AI news article illustration

PaddleOCR 3.5 brings OCR and document parsing closer to the Hugging Face ecosystem. With this release, supported PaddleOCR models can run with Hugging Face Transformers as an inference backend, so developers working inside PyTorch-based stacks can stop wiring OCR systems together by hand.

What Changed in PaddleOCR 3.5

The release introduces a flexible inference-engine interface. Developers select a backend through the engine parameter and pass backend-specific options through engine_config, meaning the switch happens without changing your existing OCR pipeline.

PaddleOCR continues to ship OCR series like PP-OCRv5 and document parsing series like PaddleOCR-VL 1.5; Transformers joins them as one of several supported runtimes.

Why It Matters for Document AI

For RAG, Document AI, and document-agent applications, the hard part often starts before the LLM. Developers must convert PDFs, scanned pages, tables, charts, formulas, and layouts into reliable structured data — and if that ingestion step is weak, downstream LLM workflows miss context or produce unreliable answers.

With PaddleOCR 3.5, this ingestion layer plugs more naturally into Transformers-centered stacks, reducing integration friction on the path from raw documents to RAG, search, analytics, and automation.

Quick Start

On a CUDA 12.6 environment the setup is a few pip installs and one command:

python -m pip install "paddleocr==3.5.0" "paddlex==3.5.2" "transformers>=5.4.0"
paddleocr ocr -i <image_url> --device gpu:0 --engine transformers

The same interface is available through the Python API, with engine_config used to tune hardware-specific options like running bfloat16 with the scaled-dot-product attention implementation.

When to Use the Transformers Backend

The Transformers backend is a natural fit for teams already relying on Hugging Face infrastructure for model loading, experimentation, and deployment, and for developers who want Hub-compatible model discovery. This release is not about replacing one backend with another — it is about giving developers the freedom to pick the runtime that fits their stack.

What This Means

For the open-source ecosystem, PaddleOCR 3.5 removes a real bottleneck: document ingestion no longer forces a separate runtime just to get text out of images and complex layouts. As more teams build agents and retrieval systems on top of OCR output, having Transformers as a first-class backend lowers the barrier between raw documents and everything downstream of them.

Frequently Asked Questions

What is new in PaddleOCR 3.5?

PaddleOCR 3.5 introduces a more flexible inference-engine interface. Developers pick a backend through the engine parameter, and Hugging Face Transformers is now a supported backend for running PP-OCRv5 and PaddleOCR-VL 1.5 models.

How do I run PaddleOCR with a Transformers backend?

Install paddleocr==3.5.0, paddlex==3.5.2, and transformers>=5.4.0, then pass engine=transformers when creating a PaddleOCR pipeline. Backend options such as dtype and device are passed through engine_config.

When should I use the Transformers backend instead of the default?

Use Transformers when your stack is already Hugging Face centered, such as RAG, search, analytics, or agent applications. If maximizing OCR or document parsing throughput is the priority, the default paddle_static backend is usually the recommended choice.

What models does PaddleOCR 3.5 support?

PaddleOCR continues to provide OCR model series such as PP-OCRv5 and document parsing series such as PaddleOCR-VL 1.5, which can now run with Transformers as an alternative inference backend while PaddleOCR manages the pipeline.

This article is based on the official announcement from huggingface.co . Read the original for full technical details.

Related Articles

Back to all news