OptimumDocumentEmbedder
Compute document embeddings using Hugging Face models accelerated by the ONNX runtime via HuggingFace Optimum. Use this component in indexing pipelines to embed documents before writing them to a document store.
Key Features
- Uses the HuggingFace Optimum library and ONNX runtime for fast, hardware-accelerated inference.
- Supports multiple execution providers including CPU (default) and TensorRT for GPU acceleration.
- Stores the computed embedding in each document's
embeddingfield, making documents ready for semantic search. - Supports optional model optimization and quantization through
optimizer_settingsandquantizer_settings. - Configurable batch size, metadata embedding, and prefix/suffix text injection.
- Compatible with any Sentence Transformers model available on Hugging Face Hub.
The embedding model you use to embed documents in your index must be the same as the embedding model you use to embed the query in your pipeline.
This means the embedders for your indexes and pipelines must match. For example, if you use CohereDocumentEmbedder to embed your documents, you should use CohereTextEmbedder with the same model to embed your queries.
Configuration
- Drag the
OptimumDocumentEmbeddercomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- On the General tab:
- Set the model field to the Hugging Face model ID you want to use, for example
sentence-transformers/all-mpnet-base-v2. - Optionally, create a secret with your Hugging Face API token and use
HF_API_TOKENas the secret key. For instructions, see Create Secrets.
- Set the model field to the Hugging Face model ID you want to use, for example
- Go to the Advanced tab to configure
onnx_execution_provider,batch_size,normalize_embeddings, and other settings.
Connections
OptimumDocumentEmbedder receives a list of documents through its documents input. It outputs the same documents with their embeddings added through its documents output.
Connect it after converters (such as TextFileToDocument or HTMLToDocument) or after DocumentSplitter to embed chunks. Connect its documents output to DocumentWriter to store embedded documents in a document store.
Source Code
To check this component's source code, open optimum_document_embedder.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
OptimumDocumentEmbedder:
type: haystack_integrations.components.embedders.optimum.optimum_document_embedder.OptimumDocumentEmbedder
init_parameters:
model: sentence-transformers/all-mpnet-base-v2
token:
type: env_var
env_vars:
- HF_API_TOKEN
strict: false
normalize_embeddings: true
onnx_execution_provider: CPUExecutionProvider
batch_size: 32
progress_bar: true
embedding_separator: "\n"
Using the Component in a Pipeline
This example shows an indexing pipeline that reads text files, splits them into chunks, embeds them using Optimum with ONNX acceleration, and writes them to an OpenSearch document store.
# haystack-pipeline
components:
TextFileToDocument:
type: haystack.components.converters.txt.TextFileToDocument
init_parameters:
encoding: utf-8
store_full_path: false
DocumentSplitter:
type: haystack.components.preprocessors.document_splitter.DocumentSplitter
init_parameters:
split_by: sentence
split_length: 100
split_overlap: 0
split_threshold: 0
splitting_function:
OptimumDocumentEmbedder:
type: haystack_integrations.components.embedders.optimum.optimum_document_embedder.OptimumDocumentEmbedder
init_parameters:
model: sentence-transformers/all-mpnet-base-v2
token:
type: env_var
env_vars:
- HF_API_TOKEN
strict: false
normalize_embeddings: true
onnx_execution_provider: CPUExecutionProvider
batch_size: 32
progress_bar: true
meta_fields_to_embed:
embedding_separator: "\n"
DocumentWriter:
type: haystack.components.writers.document_writer.DocumentWriter
init_parameters:
policy: OVERWRITE
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
index: Standard-Index-English
max_chunk_bytes: 104857600
embedding_dim: 768
return_embedding: false
method:
mappings:
settings:
create_index: true
http_auth:
use_ssl:
verify_certs:
timeout:
connections:
- sender: TextFileToDocument.documents
receiver: DocumentSplitter.documents
- sender: DocumentSplitter.documents
receiver: OptimumDocumentEmbedder.documents
- sender: OptimumDocumentEmbedder.documents
receiver: DocumentWriter.documents
max_runs_per_component: 100
metadata: {}
inputs:
files:
- TextFileToDocument.sources
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | A list of Documents to embed. |
Outputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | Documents with their embedding field populated. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
model | str | sentence-transformers/all-mpnet-base-v2 | The Hugging Face model ID to use for embedding. |
token | Optional[Secret] | Secret.from_env_var('HF_API_TOKEN', strict=False) | The Hugging Face API token used as HTTP bearer authorization. |
prefix | str | "" | A string to add to the beginning of each text before embedding. |
suffix | str | "" | A string to add to the end of each text before embedding. |
normalize_embeddings | bool | True | Whether to normalize the embeddings to unit length. |
onnx_execution_provider | str | CPUExecutionProvider | The ONNX execution provider to use for inference. Common options: CPUExecutionProvider, CUDAExecutionProvider, TensorrtExecutionProvider. |
pooling_mode | Optional[str | OptimumEmbedderPooling] | None | The pooling mode to use. When None, pooling mode is inferred from the model config. |
model_kwargs | Optional[Dict[str, Any]] | None | Additional keyword arguments to pass to the model. Overrides model, onnx_execution_provider, and token when there is a conflict. |
working_dir | Optional[str] | None | Directory for intermediate files generated during model optimization or quantization. Required when using optimizer_settings or quantizer_settings. |
optimizer_settings | Optional[OptimumEmbedderOptimizationConfig] | None | Configuration for Optimum Embedder optimization. When None, no additional optimization is applied. |
quantizer_settings | Optional[OptimumEmbedderQuantizationConfig] | None | Configuration for Optimum Embedder quantization. When None, no quantization is applied. |
batch_size | int | 32 | Number of Documents to encode at once. |
progress_bar | bool | True | Whether to show a progress bar. Disable in production to keep logs clean. |
meta_fields_to_embed | Optional[List[str]] | None | Metadata fields to embed alongside the Document text. |
embedding_separator | str | \n | Separator used to concatenate metadata fields and Document text. |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | A list of Documents to embed. |
Related Information
Was this page helpful?