Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

HuggingFaceTEISparseTextEmbedder

Embed a query string into a sparse vector using a self-hosted Hugging Face Text Embeddings Inference (TEI) server. Use this component in query pipelines to embed user queries for sparse retrieval. This component was previously called HuggingFaceAPISparseTextEmbedder.

Key Features​

  • Connects to a self-hosted Hugging Face TEI server running a sparse embedding model. HTTP calls the /embed_sparse endpoint. Set use_grpc to connect over gRPC instead.
  • Outputs a SparseEmbedding object suitable for use with sparse embedding retrievers.
  • Lightweight—no batch size or progress bar configuration needed for single-query embedding.
  • Supports both synchronous and asynchronous operation via run() and run_async().
  • No external API key is required—authentication is optional and uses a Hugging Face token if the TEI server requires it.
  • The TEI server model must match the one used in the corresponding HuggingFaceTEISparseDocumentEmbedder in the indexing pipeline.

Configuration​

  1. Drag the HuggingFaceTEISparseTextEmbedder component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    1. Set the api_base_url field to the base URL of your TEI server, for example http://localhost:8080.
    2. Optionally, create a secret with your Hugging Face token and use HF_API_TOKEN as the secret key if the TEI server requires authentication.
  4. Go to the Advanced tab to configure prefix, suffix, timeout, and use_grpc.

Connections​

HuggingFaceTEISparseTextEmbedder receives the user query as a text string through its text input, typically from the Input component. It outputs a SparseEmbedding through its sparse_embedding output, which you connect to a sparse retriever.

Source Code​

To check this component's source code, open sparse_text_embedder.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

HuggingFaceTEISparseTextEmbedder:
type: haystack_integrations.components.embedders.huggingface_api.sparse_text_embedder.HuggingFaceTEISparseTextEmbedder
init_parameters:
api_base_url: http://localhost:8080
token:
type: env_var
env_vars:
- HF_API_TOKEN
- HF_TOKEN
strict: false
timeout: 30.0

Using the Component in a Pipeline​

This example shows a query pipeline with HuggingFaceTEISparseTextEmbedder that embeds the user query into a sparse vector for retrieval.

# haystack-pipeline
components:
HuggingFaceTEISparseTextEmbedder:
type: haystack_integrations.components.embedders.huggingface_api.sparse_text_embedder.HuggingFaceTEISparseTextEmbedder
init_parameters:
api_base_url: http://localhost:8080
token:
type: env_var
env_vars:
- HF_API_TOKEN
- HF_TOKEN
strict: false
prefix: ""
suffix: ""
timeout: 30.0
sparse_retriever:
type: haystack_integrations.components.retrievers.qdrant.retriever.QdrantSparseEmbeddingRetriever
init_parameters:
document_store:
type: haystack_integrations.document_stores.qdrant.document_store.QdrantDocumentStore
init_parameters:
location: http://localhost:6333
index: default
use_sparse_embeddings: true
return_embedding: false
top_k: 10
scale_score: false

connections:
- sender: HuggingFaceTEISparseTextEmbedder.sparse_embedding
receiver: sparse_retriever.query_sparse_embedding

max_runs_per_component: 100

metadata: {}

inputs:
query:
- HuggingFaceTEISparseTextEmbedder.text

Parameters​

Inputs​

ParameterTypeDescription
textstrThe text to embed.

Outputs​

ParameterTypeDescription
sparse_embeddingSparseEmbeddingThe sparse embedding of the input text.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
api_base_urlstrhttp://localhost:8080Base URL of the TEI server, or the gRPC target when use_grpc is on.
tokenOptional[Secret]Secret.from_env_var(['HF_API_TOKEN', 'HF_TOKEN'], strict=False)Token sent as HTTP bearer authorization to the TEI server, if required.
prefixstr""A string to add before the text.
suffixstr""A string to add after the text.
timeoutOptional[float]30.0HTTP request timeout in seconds. Set to None to disable.
headersOptional[Dict[str, str]]NoneAdditional HTTP headers to send with each request.
use_grpcboolFalseConnect over gRPC instead of HTTP.

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDescription
textstrThe text to embed.