KreuzbergConverter
Convert PDFs, Office documents, images, and 75+ other file formats to Haystack Documents using the Kreuzberg document intelligence library — all locally with no external API calls.
Key Features
- Converts PDFs, DOCX, images, and 75+ formats locally without any external API dependencies.
- Supports multiple OCR backends, including Tesseract and EasyOCR, for scanned documents and images.
- Extracts per-page or chunked Documents using Kreuzberg's built-in pagination and chunking configuration.
- Detects document language automatically and enriches Document metadata with quality scores, keywords, and extraction format.
- Processes multiple files in parallel using Kreuzberg's batch extraction APIs.
- Accepts file paths, directory paths, and
ByteStreaminputs.
Configuration
- Install the
kreuzberg-haystackintegration. - Configure
KreuzbergConverterin your pipeline YAML.
Connections
KreuzbergConverter accepts file paths, Path objects, or ByteStream inputs as sources, and optional metadata as meta. It outputs a list of Document objects (documents).
Connect the pipeline's file input to the sources input. Connect the documents output to DocumentSplitter, DocumentJoiner, or DocumentWriter for further processing.
Source Code
To check this component's source code, open converter.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
kreuzberg_converter:
type: haystack_integrations.components.converters.kreuzberg.converter.KreuzbergConverter
init_parameters:
store_full_path: false
batch: true
Using the Component in a Pipeline
# haystack-pipeline
components:
kreuzberg_converter:
type: haystack_integrations.components.converters.kreuzberg.converter.KreuzbergConverter
init_parameters:
store_full_path: false
batch: true
splitter:
type: haystack.components.preprocessors.document_splitter.DocumentSplitter
init_parameters:
split_by: word
split_length: 250
split_overlap: 30
writer:
type: haystack.components.writers.document_writer.DocumentWriter
init_parameters:
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
- ${OPENSEARCH_HOST}
index: kreuzberg-demo
embedding_dim: 768
create_index: true
connections:
- sender: kreuzberg_converter.documents
receiver: splitter.documents
- sender: splitter.documents
receiver: writer.documents
inputs:
files:
- kreuzberg_converter.sources
max_runs_per_component: 100
metadata: {}
Parameters
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
config | ExtractionConfig | None | None | A kreuzberg.ExtractionConfig object to customize extraction behavior. Controls output format, OCR backend and language, per-page extraction, chunking, and more. Cannot be used together with config_path. |
config_path | str | None | None | Path to a Kreuzberg configuration file (.toml, .yaml, or .json). Cannot be used together with config. |
store_full_path | bool | False | If True, stores the full file path in Document metadata. If False, stores only the file name. |
batch | bool | True | If True, uses Kreuzberg's batch extraction APIs for parallel processing. If False, processes sources one at a time. |
easyocr_kwargs | dict | None | None | Optional keyword arguments for the EasyOCR backend, such as GPU settings and beam width. Only applied when using the easyocr OCR backend. |
Related Information
Was this page helpful?