TypeSafeDocumentClassifier
Answer typed questions about each document with a TypeSafe System One model and store the answers in the document's metadata.
Key Features
- Classifies documents with a System One decision model, such as Jev. These models don't generate text. They answer every question in one request and return calibrated probabilities.
- Supports three question types:
choice: picks one label fromcriteria, a mapping of label to description.score: rates the text on the ordered levels incriteria, a list of level descriptions.noul: returns the probability that the yes/no question ininstructionsis true.
- Stores the answers under
meta[metadata_field], keyed by question ID. - Returns documents without text to classify, and documents whose request failed, in the
failed_documentsoutput with the reason inmeta["classification_error"]. - Sends each document as its own request and runs up to
max_workersrequests at once. - Works with any server that implements the TypeSafe API, such as Ollaya, which runs open decision models locally.
Configuration
- Drag the
TypeSafeDocumentClassifiercomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- On the General tab:
- Set
questionsto the questions you want the model to answer for every document, keyed by question ID. Each question needs atype(choice,score, ornoul),instructions, and forchoiceandscorequestions,criteria. - Optionally, change
model. The default isjev-latest. - If you use a self-hosted TypeSafe-compatible server, set
api_base_urlto its URL.
- Set
- On the Authentication tab, connect your TypeSafe API key. The component reads it from the
TYPESAFE_API_KEYsecret by default. For details, see Add Secrets.
Connections
TypeSafeDocumentClassifier receives a list of documents and outputs the same documents with the answers added to their metadata. Connect its documents output to a DocumentWriter to store the enriched documents, or to a MetadataRouter to route documents by their answers. Connect failed_documents to a separate branch if you want to keep or inspect documents that couldn't be classified.
Source Code
To check this component's source code, open document_classifier.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
TypeSafeDocumentClassifier:
type: haystack_integrations.components.classifiers.typesafe.document_classifier.TypeSafeDocumentClassifier
init_parameters:
api_key:
type: env_var
env_vars:
- TYPESAFE_API_KEY
strict: false
model: jev-latest
questions:
department:
type: choice
instructions: Which department should handle this ticket?
criteria:
billing:
technical:
sales:
urgency:
type: score
instructions: How urgent is this ticket?
criteria:
- can wait
- this week
- today
refund:
type: noul
instructions: Is the customer asking for a refund?
Using the Component in an Index
This index classifies each document and writes the answers to its metadata, so you can filter on them at query time:
# haystack-pipeline
components:
TypeSafeDocumentClassifier:
type: haystack_integrations.components.classifiers.typesafe.document_classifier.TypeSafeDocumentClassifier
init_parameters:
api_key:
type: env_var
env_vars:
- TYPESAFE_API_KEY
strict: false
model: jev-latest
metadata_field: typesafe
max_workers: 3
raise_on_failure: false
questions:
department:
type: choice
instructions: Which department should handle this ticket?
criteria:
billing:
technical:
sales:
refund:
type: noul
instructions: Is the customer asking for a refund?
writer:
type: haystack.components.writers.document_writer.DocumentWriter
init_parameters:
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
index: ''
embedding_dim: 768
create_index: true
policy: OVERWRITE
connections:
- sender: TypeSafeDocumentClassifier.documents
receiver: writer.documents
max_runs_per_component: 100
metadata: {}
inputs:
documents:
- TypeSafeDocumentClassifier.documents
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | The documents to classify. |
Outputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | The classified documents. Each has a metadata_field entry mapping each question ID to its answer. |
failed_documents | List[Document] | The documents that have no text to classify or whose request failed, with the reason in meta["classification_error"]. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
questions | Dict[str, Question] | Questions to answer for every document, keyed by question ID. Each value has a type (choice, score, or noul), instructions, and for choice and score questions, criteria. | |
model | str | jev-latest | Name of the System One model to use. |
api_key | Secret | Secret.from_env_var("TYPESAFE_API_KEY") | The TypeSafe API key. Servers such as Ollaya accept any non-empty value unless they are configured with a key. |
api_base_url | Optional[str] | None | Base URL of the API. If None, the TYPESAFE_BASE_URL environment variable is used, or https://api.typesafe.ai if it is unset. |
classification_field | Optional[str] | None | Name of the document's metadata field to classify. If None, Document.content is classified. |
metadata_field | str | typesafe | Name of the metadata field the answers are written to. |
timeout | Optional[float] | None | Timeout in seconds for each request attempt. If None, the SDK default of 10 seconds is used. |
max_retries | Optional[int] | None | Maximum number of retries after a failed request attempt. If None, the SDK default of two is used. |
max_workers | int | 3 | Maximum number of concurrent requests. Must be at least one. |
raise_on_failure | bool | False | If True, the first failed request raises its error. If False, documents whose request failed are returned in failed_documents. |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Default | Description |
|---|---|---|---|
documents | List[Document] | The documents to classify. |
Related Information
Was this page helpful?