Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

PresidioTextCleaner

Remove personally identifiable information (PII) from plain text strings using Microsoft Presidio, replacing detected entities with type placeholders such as <PERSON> or <PHONE_NUMBER>.

Key Features​

  • Sanitizes raw text strings before they are sent to a language model, helping prevent PII leakage.
  • Replaces detected entities with labeled placeholders such as <PERSON> and <EMAIL_ADDRESS>.
  • Runs locally — no external API call or API key required.
  • Supports multiple languages; automatically selects the appropriate spaCy model for built-in languages.
  • Configurable confidence threshold and entity type allowlist for fine-grained control.
  • Unlike PresidioDocumentCleaner, operates on plain strings rather than Document objects, making it suitable for cleaning user queries.

Configuration​

  1. Drag the PresidioTextCleaner component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    • Set language to the ISO 639-1 code matching your text language.
    • Optionally, set entities to restrict anonymization to specific PII types.
  4. Go to the Advanced tab to adjust score_threshold or configure custom models for unsupported languages.
spaCy model dependency

PresidioTextCleaner requires a spaCy language model. For English, install en_core_web_lg with python -m spacy download en_core_web_lg. For other built-in languages, install the corresponding model listed in the spaCy model documentation. The component loads the model automatically on the first run() call.

Connections​

PresidioTextCleaner receives a list of strings, typically raw user queries from the Input component. It outputs a new list of strings with PII removed. Connect its texts output to a component that uses the cleaned text, such as a PromptBuilder or a chat generator.

Source Code​

To check this component's source code, open presidio_text_cleaner.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

PresidioTextCleaner:
type: haystack_integrations.components.preprocessors.presidio.presidio_text_cleaner.PresidioTextCleaner
init_parameters:
language: en
entities:
score_threshold: 0.35
models:

Using the Component in a Pipeline​

# haystack-pipeline
components:
PresidioTextCleaner:
type: haystack_integrations.components.preprocessors.presidio.presidio_text_cleaner.PresidioTextCleaner
init_parameters:
language: en
entities:
- PERSON
- PHONE_NUMBER
- EMAIL_ADDRESS
score_threshold: 0.35
models:

connections: []

max_runs_per_component: 100

metadata: {}

inputs:
texts:
- PresidioTextCleaner.texts

outputs:
texts: PresidioTextCleaner.texts

Parameters​

Inputs​

ParameterTypeDescription
textsList[str]List of strings to anonymize.

Outputs​

ParameterTypeDescription
textsList[str]Cleaned strings with PII replaced by entity type placeholders.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
languagestr"en"ISO 639-1 language code for PII detection. For built-in languages such as "de", "fr", and "es", the appropriate spaCy model is loaded automatically. For unsupported languages, use models to specify a custom model. See Presidio supported languages.
entitiesOptional[List[str]]NoneList of PII entity types to detect and anonymize, for example ["PERSON", "PHONE_NUMBER"]. If None, all supported entity types are used. See Presidio supported entities.
score_thresholdfloat0.35Minimum confidence score (0–1) for a detected entity to be anonymized.
modelsOptional[List[Dict[str, str]]]NoneAdvanced override: list of spaCy model configurations. Each entry must contain "lang_code" and "model_name" keys, for example [{"lang_code": "fr", "model_name": "fr_core_news_md"}]. Use this only when you need a specific model variant or a language not covered by the built-in mapping.