Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

PresidioEntityExtractor

Detect personally identifiable information (PII) entities in Haystack Document objects using Microsoft Presidio, storing the detected entities in each document's metadata for downstream processing.

Key Features​

  • Identifies PII entities in document text content and stores their type, character offsets, and confidence score in document metadata under the key "entities".
  • Does not modify or remove document text — the original content is preserved; only metadata is updated.
  • Runs locally — no external API call or API key required.
  • Supports multiple languages; automatically selects the appropriate spaCy model for built-in languages.
  • Documents without text content pass through unchanged.
  • Configurable confidence threshold and entity type allowlist for precise entity detection.

Configuration​

  1. Drag the PresidioEntityExtractor component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    • Set language to the ISO 639-1 code of your document language.
    • Optionally, set entities to restrict detection to specific PII types such as PERSON or EMAIL_ADDRESS.
  4. Go to the Advanced tab to adjust score_threshold or configure custom models for unsupported languages.
spaCy model dependency

PresidioEntityExtractor requires a spaCy language model. For English, install en_core_web_lg with python -m spacy download en_core_web_lg. For other built-in languages, install the corresponding model listed in the spaCy model documentation. The component loads the model automatically on the first run() call.

Connections​

PresidioEntityExtractor receives a list of documents, typically from a document converter or retriever. It outputs the same documents with detected entities added to each document's metadata. Connect its documents output to a document writer, PresidioDocumentCleaner, or any other component that processes document metadata.

Source Code​

To check this component's source code, open presidio_entity_extractor.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

PresidioEntityExtractor:
type: haystack_integrations.components.extractors.presidio.presidio_entity_extractor.PresidioEntityExtractor
init_parameters:
language: en
entities:
score_threshold: 0.35
models:

Using the Component in a Pipeline​

# haystack-pipeline
components:
PresidioEntityExtractor:
type: haystack_integrations.components.extractors.presidio.presidio_entity_extractor.PresidioEntityExtractor
init_parameters:
language: en
entities:
- PERSON
- EMAIL_ADDRESS
- PHONE_NUMBER
score_threshold: 0.5
models:

connections: []

max_runs_per_component: 100

metadata: {}

inputs:
documents:
- PresidioEntityExtractor.documents

outputs:
documents: PresidioEntityExtractor.documents

Parameters​

Inputs​

ParameterTypeDescription
documentsList[Document]Documents to analyze for PII entities.

Outputs​

ParameterTypeDescription
documentsList[Document]Documents with detected PII entities stored in metadata under the key "entities". Each entry contains entity_type, start, end (character offsets), and score (confidence). Documents without text content are passed through unchanged.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
languagestr"en"ISO 639-1 language code for PII detection. For built-in languages such as "de", "fr", and "es", the appropriate spaCy model is loaded automatically. For unsupported languages, use models to specify a custom model. See Presidio supported languages.
entitiesOptional[List[str]]NoneList of PII entity types to detect, for example ["PERSON", "EMAIL_ADDRESS"]. If None, all supported entity types are detected. See Presidio supported entities.
score_thresholdfloat0.35Minimum confidence score (0–1) for a detected entity to be included.
modelsOptional[List[Dict[str, str]]]NoneAdvanced override: list of spaCy model configurations. Each entry must contain "lang_code" and "model_name" keys, for example [{"lang_code": "fr", "model_name": "fr_core_news_md"}]. Use this only when you need a specific model variant or a language not covered by the built-in mapping.