PresidioTextCleaner
Remove personally identifiable information (PII) from plain text strings using Microsoft Presidio, replacing detected entities with type placeholders such as <PERSON> or <PHONE_NUMBER>.
Key Features
- Sanitizes raw text strings before they are sent to a language model, helping prevent PII leakage.
- Replaces detected entities with labeled placeholders such as
<PERSON>and<EMAIL_ADDRESS>. - Runs locally — no external API call or API key required.
- Supports multiple languages; automatically selects the appropriate spaCy model for built-in languages.
- Configurable confidence threshold and entity type allowlist for fine-grained control.
- Unlike
PresidioDocumentCleaner, operates on plain strings rather thanDocumentobjects, making it suitable for cleaning user queries.
Configuration
- Drag the
PresidioTextCleanercomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- On the General tab:
- Set
languageto the ISO 639-1 code matching your text language. - Optionally, set
entitiesto restrict anonymization to specific PII types.
- Set
- Go to the Advanced tab to adjust
score_thresholdor configure custommodelsfor unsupported languages.
PresidioTextCleaner requires a spaCy language model. For English, install en_core_web_lg with python -m spacy download en_core_web_lg. For other built-in languages, install the corresponding model listed in the spaCy model documentation. The component loads the model automatically on the first run() call.
Connections
PresidioTextCleaner receives a list of strings, typically raw user queries from the Input component. It outputs a new list of strings with PII removed. Connect its texts output to a component that uses the cleaned text, such as a PromptBuilder or a chat generator.
Source Code
To check this component's source code, open presidio_text_cleaner.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
PresidioTextCleaner:
type: haystack_integrations.components.preprocessors.presidio.presidio_text_cleaner.PresidioTextCleaner
init_parameters:
language: en
entities:
score_threshold: 0.35
models:
Using the Component in a Pipeline
# haystack-pipeline
components:
PresidioTextCleaner:
type: haystack_integrations.components.preprocessors.presidio.presidio_text_cleaner.PresidioTextCleaner
init_parameters:
language: en
entities:
- PERSON
- PHONE_NUMBER
- EMAIL_ADDRESS
score_threshold: 0.35
models:
connections: []
max_runs_per_component: 100
metadata: {}
inputs:
texts:
- PresidioTextCleaner.texts
outputs:
texts: PresidioTextCleaner.texts
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
texts | List[str] | List of strings to anonymize. |
Outputs
| Parameter | Type | Description |
|---|---|---|
texts | List[str] | Cleaned strings with PII replaced by entity type placeholders. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
language | str | "en" | ISO 639-1 language code for PII detection. For built-in languages such as "de", "fr", and "es", the appropriate spaCy model is loaded automatically. For unsupported languages, use models to specify a custom model. See Presidio supported languages. |
entities | Optional[List[str]] | None | List of PII entity types to detect and anonymize, for example ["PERSON", "PHONE_NUMBER"]. If None, all supported entity types are used. See Presidio supported entities. |
score_threshold | float | 0.35 | Minimum confidence score (0–1) for a detected entity to be anonymized. |
models | Optional[List[Dict[str, str]]] | None | Advanced override: list of spaCy model configurations. Each entry must contain "lang_code" and "model_name" keys, for example [{"lang_code": "fr", "model_name": "fr_core_news_md"}]. Use this only when you need a specific model variant or a language not covered by the built-in mapping. |
Related Information
Was this page helpful?