HuggingFaceTEISparseTextEmbedder
Embed a query string into a sparse vector using a self-hosted Hugging Face Text Embeddings Inference (TEI) server. Use this component in query pipelines to embed user queries for sparse retrieval. This component was previously called HuggingFaceAPISparseTextEmbedder.
Key Features
- Connects to a self-hosted Hugging Face TEI server running a sparse embedding model. HTTP calls the
/embed_sparseendpoint. Setuse_grpcto connect over gRPC instead. - Outputs a
SparseEmbeddingobject suitable for use with sparse embedding retrievers. - Lightweight—no batch size or progress bar configuration needed for single-query embedding.
- Supports both synchronous and asynchronous operation via
run()andrun_async(). - No external API key is required—authentication is optional and uses a Hugging Face token if the TEI server requires it.
- The TEI server model must match the one used in the corresponding
HuggingFaceTEISparseDocumentEmbedderin the indexing pipeline.
Configuration
- Drag the
HuggingFaceTEISparseTextEmbeddercomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- On the General tab:
- Set the api_base_url field to the base URL of your TEI server, for example
http://localhost:8080. - Optionally, create a secret with your Hugging Face token and use
HF_API_TOKENas the secret key if the TEI server requires authentication.
- Set the api_base_url field to the base URL of your TEI server, for example
- Go to the Advanced tab to configure
prefix,suffix,timeout, anduse_grpc.
Connections
HuggingFaceTEISparseTextEmbedder receives the user query as a text string through its text input, typically from the Input component. It outputs a SparseEmbedding through its sparse_embedding output, which you connect to a sparse retriever.
Source Code
To check this component's source code, open sparse_text_embedder.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
HuggingFaceTEISparseTextEmbedder:
type: haystack_integrations.components.embedders.huggingface_api.sparse_text_embedder.HuggingFaceTEISparseTextEmbedder
init_parameters:
api_base_url: http://localhost:8080
token:
type: env_var
env_vars:
- HF_API_TOKEN
- HF_TOKEN
strict: false
timeout: 30.0
Using the Component in a Pipeline
This example shows a query pipeline with HuggingFaceTEISparseTextEmbedder that embeds the user query into a sparse vector for retrieval.
# haystack-pipeline
components:
HuggingFaceTEISparseTextEmbedder:
type: haystack_integrations.components.embedders.huggingface_api.sparse_text_embedder.HuggingFaceTEISparseTextEmbedder
init_parameters:
api_base_url: http://localhost:8080
token:
type: env_var
env_vars:
- HF_API_TOKEN
- HF_TOKEN
strict: false
prefix: ""
suffix: ""
timeout: 30.0
sparse_retriever:
type: haystack_integrations.components.retrievers.qdrant.retriever.QdrantSparseEmbeddingRetriever
init_parameters:
document_store:
type: haystack_integrations.document_stores.qdrant.document_store.QdrantDocumentStore
init_parameters:
location: http://localhost:6333
index: default
use_sparse_embeddings: true
return_embedding: false
top_k: 10
scale_score: false
connections:
- sender: HuggingFaceTEISparseTextEmbedder.sparse_embedding
receiver: sparse_retriever.query_sparse_embedding
max_runs_per_component: 100
metadata: {}
inputs:
query:
- HuggingFaceTEISparseTextEmbedder.text
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
text | str | The text to embed. |
Outputs
| Parameter | Type | Description |
|---|---|---|
sparse_embedding | SparseEmbedding | The sparse embedding of the input text. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
api_base_url | str | http://localhost:8080 | Base URL of the TEI server, or the gRPC target when use_grpc is on. |
token | Optional[Secret] | Secret.from_env_var(['HF_API_TOKEN', 'HF_TOKEN'], strict=False) | Token sent as HTTP bearer authorization to the TEI server, if required. |
prefix | str | "" | A string to add before the text. |
suffix | str | "" | A string to add after the text. |
timeout | Optional[float] | 30.0 | HTTP request timeout in seconds. Set to None to disable. |
headers | Optional[Dict[str, str]] | None | Additional HTTP headers to send with each request. |
use_grpc | bool | False | Connect over gRPC instead of HTTP. |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Description |
|---|---|---|
text | str | The text to embed. |
Related Information
Was this page helpful?