PresidioEntityExtractor
Detect personally identifiable information (PII) entities in Haystack Document objects using Microsoft Presidio, storing the detected entities in each document's metadata for downstream processing.
Key Features
- Identifies PII entities in document text content and stores their type, character offsets, and confidence score in document metadata under the key
"entities". - Does not modify or remove document text — the original content is preserved; only metadata is updated.
- Runs locally — no external API call or API key required.
- Supports multiple languages; automatically selects the appropriate spaCy model for built-in languages.
- Documents without text content pass through unchanged.
- Configurable confidence threshold and entity type allowlist for precise entity detection.
Configuration
- Drag the
PresidioEntityExtractorcomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- On the General tab:
- Set
languageto the ISO 639-1 code of your document language. - Optionally, set
entitiesto restrict detection to specific PII types such asPERSONorEMAIL_ADDRESS.
- Set
- Go to the Advanced tab to adjust
score_thresholdor configure custommodelsfor unsupported languages.
PresidioEntityExtractor requires a spaCy language model. For English, install en_core_web_lg with python -m spacy download en_core_web_lg. For other built-in languages, install the corresponding model listed in the spaCy model documentation. The component loads the model automatically on the first run() call.
Connections
PresidioEntityExtractor receives a list of documents, typically from a document converter or retriever. It outputs the same documents with detected entities added to each document's metadata. Connect its documents output to a document writer, PresidioDocumentCleaner, or any other component that processes document metadata.
Source Code
To check this component's source code, open presidio_entity_extractor.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
PresidioEntityExtractor:
type: haystack_integrations.components.extractors.presidio.presidio_entity_extractor.PresidioEntityExtractor
init_parameters:
language: en
entities:
score_threshold: 0.35
models:
Using the Component in a Pipeline
# haystack-pipeline
components:
PresidioEntityExtractor:
type: haystack_integrations.components.extractors.presidio.presidio_entity_extractor.PresidioEntityExtractor
init_parameters:
language: en
entities:
- PERSON
- EMAIL_ADDRESS
- PHONE_NUMBER
score_threshold: 0.5
models:
connections: []
max_runs_per_component: 100
metadata: {}
inputs:
documents:
- PresidioEntityExtractor.documents
outputs:
documents: PresidioEntityExtractor.documents
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | Documents to analyze for PII entities. |
Outputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | Documents with detected PII entities stored in metadata under the key "entities". Each entry contains entity_type, start, end (character offsets), and score (confidence). Documents without text content are passed through unchanged. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
language | str | "en" | ISO 639-1 language code for PII detection. For built-in languages such as "de", "fr", and "es", the appropriate spaCy model is loaded automatically. For unsupported languages, use models to specify a custom model. See Presidio supported languages. |
entities | Optional[List[str]] | None | List of PII entity types to detect, for example ["PERSON", "EMAIL_ADDRESS"]. If None, all supported entity types are detected. See Presidio supported entities. |
score_threshold | float | 0.35 | Minimum confidence score (0–1) for a detected entity to be included. |
models | Optional[List[Dict[str, str]]] | None | Advanced override: list of spaCy model configurations. Each entry must contain "lang_code" and "model_name" keys, for example [{"lang_code": "fr", "model_name": "fr_core_news_md"}]. Use this only when you need a specific model variant or a language not covered by the built-in mapping. |
Related Information
Was this page helpful?