Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

TypeSafeDocumentClassifier

Answer typed questions about each document with a TypeSafe System One model and store the answers in the document's metadata.

Key Features​

  • Classifies documents with a System One decision model, such as Jev. These models don't generate text. They answer every question in one request and return calibrated probabilities.
  • Supports three question types:
    • choice: picks one label from criteria, a mapping of label to description.
    • score: rates the text on the ordered levels in criteria, a list of level descriptions.
    • noul: returns the probability that the yes/no question in instructions is true.
  • Stores the answers under meta[metadata_field], keyed by question ID.
  • Returns documents without text to classify, and documents whose request failed, in the failed_documents output with the reason in meta["classification_error"].
  • Sends each document as its own request and runs up to max_workers requests at once.
  • Works with any server that implements the TypeSafe API, such as Ollaya, which runs open decision models locally.

Configuration​

  1. Drag the TypeSafeDocumentClassifier component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    • Set questions to the questions you want the model to answer for every document, keyed by question ID. Each question needs a type (choice, score, or noul), instructions, and for choice and score questions, criteria.
    • Optionally, change model. The default is jev-latest.
    • If you use a self-hosted TypeSafe-compatible server, set api_base_url to its URL.
  4. On the Authentication tab, connect your TypeSafe API key. The component reads it from the TYPESAFE_API_KEY secret by default. For details, see Add Secrets.

Connections​

TypeSafeDocumentClassifier receives a list of documents and outputs the same documents with the answers added to their metadata. Connect its documents output to a DocumentWriter to store the enriched documents, or to a MetadataRouter to route documents by their answers. Connect failed_documents to a separate branch if you want to keep or inspect documents that couldn't be classified.

Source Code​

To check this component's source code, open document_classifier.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

TypeSafeDocumentClassifier:
type: haystack_integrations.components.classifiers.typesafe.document_classifier.TypeSafeDocumentClassifier
init_parameters:
api_key:
type: env_var
env_vars:
- TYPESAFE_API_KEY
strict: false
model: jev-latest
questions:
department:
type: choice
instructions: Which department should handle this ticket?
criteria:
billing:
technical:
sales:
urgency:
type: score
instructions: How urgent is this ticket?
criteria:
- can wait
- this week
- today
refund:
type: noul
instructions: Is the customer asking for a refund?

Using the Component in an Index​

This index classifies each document and writes the answers to its metadata, so you can filter on them at query time:

# haystack-pipeline
components:
TypeSafeDocumentClassifier:
type: haystack_integrations.components.classifiers.typesafe.document_classifier.TypeSafeDocumentClassifier
init_parameters:
api_key:
type: env_var
env_vars:
- TYPESAFE_API_KEY
strict: false
model: jev-latest
metadata_field: typesafe
max_workers: 3
raise_on_failure: false
questions:
department:
type: choice
instructions: Which department should handle this ticket?
criteria:
billing:
technical:
sales:
refund:
type: noul
instructions: Is the customer asking for a refund?

writer:
type: haystack.components.writers.document_writer.DocumentWriter
init_parameters:
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
index: ''
embedding_dim: 768
create_index: true
policy: OVERWRITE

connections:
- sender: TypeSafeDocumentClassifier.documents
receiver: writer.documents

max_runs_per_component: 100

metadata: {}

inputs:
documents:
- TypeSafeDocumentClassifier.documents

Parameters​

Inputs​

ParameterTypeDescription
documentsList[Document]The documents to classify.

Outputs​

ParameterTypeDescription
documentsList[Document]The classified documents. Each has a metadata_field entry mapping each question ID to its answer.
failed_documentsList[Document]The documents that have no text to classify or whose request failed, with the reason in meta["classification_error"].

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
questionsDict[str, Question]Questions to answer for every document, keyed by question ID. Each value has a type (choice, score, or noul), instructions, and for choice and score questions, criteria.
modelstrjev-latestName of the System One model to use.
api_keySecretSecret.from_env_var("TYPESAFE_API_KEY")The TypeSafe API key. Servers such as Ollaya accept any non-empty value unless they are configured with a key.
api_base_urlOptional[str]NoneBase URL of the API. If None, the TYPESAFE_BASE_URL environment variable is used, or https://api.typesafe.ai if it is unset.
classification_fieldOptional[str]NoneName of the document's metadata field to classify. If None, Document.content is classified.
metadata_fieldstrtypesafeName of the metadata field the answers are written to.
timeoutOptional[float]NoneTimeout in seconds for each request attempt. If None, the SDK default of 10 seconds is used.
max_retriesOptional[int]NoneMaximum number of retries after a failed request attempt. If None, the SDK default of two is used.
max_workersint3Maximum number of concurrent requests. Must be at least one.
raise_on_failureboolFalseIf True, the first failed request raises its error. If False, documents whose request failed are returned in failed_documents.

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDefaultDescription
documentsList[Document]The documents to classify.