Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

PaddleOCRVLDocumentConverter

Extract text from PDF files and images using PaddleOCR's vision-language document parsing API.

Key Features​

  • Extracts text and structured content from PDFs and images using PaddleOCR's VL model.
  • Supports layout detection, chart recognition, seal recognition, and table extraction.
  • Auto-detects file type from file extension or MIME type.
  • Returns both Haystack Document objects and raw PaddleOCR API responses.
  • Supports optional document orientation classification and image unwarping for scanned documents.
  • Configurable VLM sampling parameters such as temperature, top_p, and max_new_tokens.

Configuration​

  1. Sign in to AI Studio and get an access token.
  2. Deploy a PaddleOCR model endpoint and note the base URL.
  3. Set the PADDLEOCR_ACCESS_TOKEN environment variable. For instructions, see Create Secrets.
  4. Configure PaddleOCRVLDocumentConverter in your pipeline YAML with the base_url of your deployed endpoint.

Connections​

PaddleOCRVLDocumentConverter accepts file paths, Path objects, or ByteStream inputs as sources, and optional metadata as meta. It outputs a list of Document objects (documents) containing extracted text as Markdown with page breaks, and a list of raw API responses (raw_paddleocr_responses).

Connect the pipeline's file input to the sources input. Connect the documents output to DocumentSplitter or DocumentWriter for further processing.

Source Code​

To check this component's source code, open paddleocr_vl_document_converter.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

paddleocr_converter:
type: haystack_integrations.components.converters.paddleocr.paddleocr_vl_document_converter.PaddleOCRVLDocumentConverter
init_parameters:
base_url: http://your-paddleocr-endpoint.aistudio-app.com
access_token:
type: env_var
env_vars:
- PADDLEOCR_ACCESS_TOKEN
strict: false

Using the Component in a Pipeline​

# haystack-pipeline
components:
paddleocr_converter:
type: haystack_integrations.components.converters.paddleocr.paddleocr_vl_document_converter.PaddleOCRVLDocumentConverter
init_parameters:
base_url: http://your-paddleocr-endpoint.aistudio-app.com
access_token:
type: env_var
env_vars:
- PADDLEOCR_ACCESS_TOKEN
strict: false
splitter:
type: haystack.components.preprocessors.document_splitter.DocumentSplitter
init_parameters:
split_by: word
split_length: 250
split_overlap: 30
writer:
type: haystack.components.writers.document_writer.DocumentWriter
init_parameters:
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
- ${OPENSEARCH_HOST}
index: paddleocr-demo
embedding_dim: 768
create_index: true

connections:
- sender: paddleocr_converter.documents
receiver: splitter.documents
- sender: splitter.documents
receiver: writer.documents

inputs:
files:
- paddleocr_converter.sources

max_runs_per_component: 100

metadata: {}

Parameters​

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
base_urlstr | NoneNoneBase URL of the PaddleOCR API endpoint. Falls back to the PADDLEOCR_BASE_URL environment variable if not set.
access_tokenSecretSecret.from_env_var(["PADDLEOCR_ACCESS_TOKEN", "AISTUDIO_ACCESS_TOKEN"])Access token for the PaddleOCR API.
modelstrModel.PADDLE_OCR_VL_16Document parsing model to use.
file_typestr | NoneNoneInput file type: pdf, image, or None for auto-detection from file extension or MIME type.
use_doc_orientation_classifybool | NoneFalseEnable document orientation classification for rotated scans.
use_doc_unwarpingbool | NoneFalseEnable text image unwarping for curved or bent documents.
use_layout_detectionbool | NoneNoneEnable layout detection.
use_chart_recognitionbool | NoneNoneEnable chart recognition.
use_seal_recognitionbool | NoneNoneEnable seal recognition.
use_ocr_for_image_blockbool | NoneNoneRecognize text inside image blocks within documents.
temperaturefloat | NoneNoneTemperature for VLM sampling.
max_new_tokensint | NoneNoneMaximum number of tokens to generate.
merge_tablesbool | NoneNoneMerge tables across pages.
additional_paramsdict | NoneNoneExtra options passed to PaddleOCRVLOptions.extra_options.