Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

OptimumTextEmbedder

Embed a query string using Hugging Face models accelerated by the ONNX runtime via HuggingFace Optimum. Use this component in query pipelines to transform user queries into vectors for embedding-based retrieval.

Key Features​

  • Uses the HuggingFace Optimum library and ONNX runtime for fast, hardware-accelerated text embedding.
  • Outputs a float vector embedding suitable for use with embedding retrievers.
  • Supports multiple execution providers including CPU (default) and TensorRT for GPU acceleration.
  • Compatible with any Sentence Transformers model available on Hugging Face Hub.
  • Supports optional model optimization and quantization through optimizer_settings and quantizer_settings.
  • The model must match the one used in the corresponding OptimumDocumentEmbedder in the indexing pipeline.

Configuration​

  1. Drag the OptimumTextEmbedder component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    1. Set the model field to the same Hugging Face model ID you used for indexing, for example sentence-transformers/all-mpnet-base-v2.
    2. Optionally, create a secret with your Hugging Face API token and use HF_API_TOKEN as the secret key. For instructions, see Create Secrets.
  4. Go to the Advanced tab to configure onnx_execution_provider, normalize_embeddings, and other settings.

Connections​

OptimumTextEmbedder receives the user query as a text string through its text input, typically from the Input component. It outputs a float vector through its embedding output, which you connect to an embedding retriever such as OpenSearchEmbeddingRetriever.

Source Code​

To check this component's source code, open optimum_text_embedder.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

OptimumTextEmbedder:
type: haystack_integrations.components.embedders.optimum.optimum_text_embedder.OptimumTextEmbedder
init_parameters:
model: sentence-transformers/all-mpnet-base-v2
token:
type: env_var
env_vars:
- HF_API_TOKEN
strict: false
normalize_embeddings: true
onnx_execution_provider: CPUExecutionProvider

Using the Component in a Pipeline​

This example shows a query pipeline with OptimumTextEmbedder that embeds the user query and sends it to OpenSearchEmbeddingRetriever to find matching documents.

# haystack-pipeline
components:
OptimumTextEmbedder:
type: haystack_integrations.components.embedders.optimum.optimum_text_embedder.OptimumTextEmbedder
init_parameters:
model: sentence-transformers/all-mpnet-base-v2
token:
type: env_var
env_vars:
- HF_API_TOKEN
strict: false
normalize_embeddings: true
onnx_execution_provider: CPUExecutionProvider
prefix: ""
suffix: ""
OpenSearchEmbeddingRetriever:
type: haystack_integrations.components.retrievers.opensearch.embedding_retriever.OpenSearchEmbeddingRetriever
init_parameters:
filters:
top_k: 10
filter_policy: replace
custom_query:
raise_on_failure: true
efficient_filtering: true
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
index: Standard-Index-English
max_chunk_bytes: 104857600
embedding_dim: 768
return_embedding: false
method:
mappings:
settings:
create_index: true
http_auth:
use_ssl:
verify_certs:
timeout:

connections:
- sender: OptimumTextEmbedder.embedding
receiver: OpenSearchEmbeddingRetriever.query_embedding

max_runs_per_component: 100

metadata: {}

inputs:
query:
- OptimumTextEmbedder.text

Parameters​

Inputs​

ParameterTypeDescription
textstrThe text to embed.

Outputs​

ParameterTypeDescription
embeddingList[float]The embedding of the input text.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
modelstrsentence-transformers/all-mpnet-base-v2The Hugging Face model ID to use for embedding.
tokenOptional[Secret]Secret.from_env_var('HF_API_TOKEN', strict=False)The Hugging Face API token used as HTTP bearer authorization.
prefixstr""A string to add to the beginning of the text before embedding.
suffixstr""A string to add to the end of the text before embedding.
normalize_embeddingsboolTrueWhether to normalize the embeddings to unit length.
onnx_execution_providerstrCPUExecutionProviderThe ONNX execution provider to use for inference. Common options: CPUExecutionProvider, CUDAExecutionProvider, TensorrtExecutionProvider.
pooling_modeOptional[str | OptimumEmbedderPooling]NoneThe pooling mode to use. When None, pooling mode is inferred from the model config.
model_kwargsOptional[Dict[str, Any]]NoneAdditional keyword arguments to pass to the model. Overrides model, onnx_execution_provider, and token when there is a conflict.
working_dirOptional[str]NoneDirectory for intermediate files generated during model optimization or quantization. Required when using optimizer_settings or quantizer_settings.
optimizer_settingsOptional[OptimumEmbedderOptimizationConfig]NoneConfiguration for Optimum Embedder optimization. When None, no additional optimization is applied.
quantizer_settingsOptional[OptimumEmbedderQuantizationConfig]NoneConfiguration for Optimum Embedder quantization. When None, no quantization is applied.

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDescription
textstrThe text to embed.