Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

KreuzbergConverter

Convert PDFs, Office documents, images, and 75+ other file formats to Haystack Documents using the Kreuzberg document intelligence library — all locally with no external API calls.

Key Features​

  • Converts PDFs, DOCX, images, and 75+ formats locally without any external API dependencies.
  • Supports multiple OCR backends, including Tesseract and EasyOCR, for scanned documents and images.
  • Extracts per-page or chunked Documents using Kreuzberg's built-in pagination and chunking configuration.
  • Detects document language automatically and enriches Document metadata with quality scores, keywords, and extraction format.
  • Processes multiple files in parallel using Kreuzberg's batch extraction APIs.
  • Accepts file paths, directory paths, and ByteStream inputs.

Configuration​

  1. Install the kreuzberg-haystack integration.
  2. Configure KreuzbergConverter in your pipeline YAML.

Connections​

KreuzbergConverter accepts file paths, Path objects, or ByteStream inputs as sources, and optional metadata as meta. It outputs a list of Document objects (documents).

Connect the pipeline's file input to the sources input. Connect the documents output to DocumentSplitter, DocumentJoiner, or DocumentWriter for further processing.

Source Code​

To check this component's source code, open converter.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

kreuzberg_converter:
type: haystack_integrations.components.converters.kreuzberg.converter.KreuzbergConverter
init_parameters:
store_full_path: false
batch: true

Using the Component in a Pipeline​

# haystack-pipeline
components:
kreuzberg_converter:
type: haystack_integrations.components.converters.kreuzberg.converter.KreuzbergConverter
init_parameters:
store_full_path: false
batch: true
splitter:
type: haystack.components.preprocessors.document_splitter.DocumentSplitter
init_parameters:
split_by: word
split_length: 250
split_overlap: 30
writer:
type: haystack.components.writers.document_writer.DocumentWriter
init_parameters:
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
- ${OPENSEARCH_HOST}
index: kreuzberg-demo
embedding_dim: 768
create_index: true

connections:
- sender: kreuzberg_converter.documents
receiver: splitter.documents
- sender: splitter.documents
receiver: writer.documents

inputs:
files:
- kreuzberg_converter.sources

max_runs_per_component: 100

metadata: {}

Parameters​

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
configExtractionConfig | NoneNoneA kreuzberg.ExtractionConfig object to customize extraction behavior. Controls output format, OCR backend and language, per-page extraction, chunking, and more. Cannot be used together with config_path.
config_pathstr | NoneNonePath to a Kreuzberg configuration file (.toml, .yaml, or .json). Cannot be used together with config.
store_full_pathboolFalseIf True, stores the full file path in Document metadata. If False, stores only the file name.
batchboolTrueIf True, uses Kreuzberg's batch extraction APIs for parallel processing. If False, processes sources one at a time.
easyocr_kwargsdict | NoneNoneOptional keyword arguments for the EasyOCR backend, such as GPU settings and beam width. Only applied when using the easyocr OCR backend.