PaddleOCRVLDocumentConverter
Extract text from PDF files and images using PaddleOCR's vision-language document parsing API.
Key Features
- Extracts text and structured content from PDFs and images using PaddleOCR's VL model.
- Supports layout detection, chart recognition, seal recognition, and table extraction.
- Auto-detects file type from file extension or MIME type.
- Returns both Haystack
Documentobjects and raw PaddleOCR API responses. - Supports optional document orientation classification and image unwarping for scanned documents.
- Configurable VLM sampling parameters such as temperature,
top_p, andmax_new_tokens.
Configuration
- Sign in to AI Studio and get an access token.
- Deploy a PaddleOCR model endpoint and note the base URL.
- Set the
PADDLEOCR_ACCESS_TOKENenvironment variable. For instructions, see Create Secrets. - Configure
PaddleOCRVLDocumentConverterin your pipeline YAML with thebase_urlof your deployed endpoint.
Connections
PaddleOCRVLDocumentConverter accepts file paths, Path objects, or ByteStream inputs as sources, and optional metadata as meta. It outputs a list of Document objects (documents) containing extracted text as Markdown with page breaks, and a list of raw API responses (raw_paddleocr_responses).
Connect the pipeline's file input to the sources input. Connect the documents output to DocumentSplitter or DocumentWriter for further processing.
Source Code
To check this component's source code, open paddleocr_vl_document_converter.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
paddleocr_converter:
type: haystack_integrations.components.converters.paddleocr.paddleocr_vl_document_converter.PaddleOCRVLDocumentConverter
init_parameters:
base_url: http://your-paddleocr-endpoint.aistudio-app.com
access_token:
type: env_var
env_vars:
- PADDLEOCR_ACCESS_TOKEN
strict: false
Using the Component in a Pipeline
# haystack-pipeline
components:
paddleocr_converter:
type: haystack_integrations.components.converters.paddleocr.paddleocr_vl_document_converter.PaddleOCRVLDocumentConverter
init_parameters:
base_url: http://your-paddleocr-endpoint.aistudio-app.com
access_token:
type: env_var
env_vars:
- PADDLEOCR_ACCESS_TOKEN
strict: false
splitter:
type: haystack.components.preprocessors.document_splitter.DocumentSplitter
init_parameters:
split_by: word
split_length: 250
split_overlap: 30
writer:
type: haystack.components.writers.document_writer.DocumentWriter
init_parameters:
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
- ${OPENSEARCH_HOST}
index: paddleocr-demo
embedding_dim: 768
create_index: true
connections:
- sender: paddleocr_converter.documents
receiver: splitter.documents
- sender: splitter.documents
receiver: writer.documents
inputs:
files:
- paddleocr_converter.sources
max_runs_per_component: 100
metadata: {}
Parameters
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
base_url | str | None | None | Base URL of the PaddleOCR API endpoint. Falls back to the PADDLEOCR_BASE_URL environment variable if not set. |
access_token | Secret | Secret.from_env_var(["PADDLEOCR_ACCESS_TOKEN", "AISTUDIO_ACCESS_TOKEN"]) | Access token for the PaddleOCR API. |
model | str | Model.PADDLE_OCR_VL_16 | Document parsing model to use. |
file_type | str | None | None | Input file type: pdf, image, or None for auto-detection from file extension or MIME type. |
use_doc_orientation_classify | bool | None | False | Enable document orientation classification for rotated scans. |
use_doc_unwarping | bool | None | False | Enable text image unwarping for curved or bent documents. |
use_layout_detection | bool | None | None | Enable layout detection. |
use_chart_recognition | bool | None | None | Enable chart recognition. |
use_seal_recognition | bool | None | None | Enable seal recognition. |
use_ocr_for_image_block | bool | None | None | Recognize text inside image blocks within documents. |
temperature | float | None | None | Temperature for VLM sampling. |
max_new_tokens | int | None | None | Maximum number of tokens to generate. |
merge_tables | bool | None | None | Merge tables across pages. |
additional_params | dict | None | None | Extra options passed to PaddleOCRVLOptions.extra_options. |
Related Information
Was this page helpful?