ElasticsearchSparseEmbeddingRetriever
Retrieve documents from an ElasticsearchDocumentStore using sparse vector similarity search with client-side sparse embeddings such as ELSER.
Key Features
- Retrieves documents using sparse vector similarity against a pre-indexed
sparse_vector_field. - Works with any client-side sparse embedding model, including ELSER.
- Configurable number of results with
top_k. - Supports metadata filtering to narrow down the search space.
- Configurable filter policy (
replaceormerge) for runtime filters. - Supports both synchronous and asynchronous execution.
Configuration
- Drag the
ElasticsearchSparseEmbeddingRetrievercomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- On the General tab:
- Configure the
document_storeconnection. Make sure theElasticsearchDocumentStorehassparse_vector_fieldset to the name of the field where your sparse embeddings are stored. - Set
top_kto the maximum number of documents to retrieve.
- Configure the
- Go to the Advanced tab to configure
filtersandfilter_policy.
This retriever accepts a pre-computed SparseEmbedding object as its query input. Use it together with a sparse text embedder such as ELSER to generate the query embedding before retrieval. For server-side inference-based sparse retrieval, use ElasticsearchInferenceSparseRetriever instead.
Connections
ElasticsearchSparseEmbeddingRetriever receives a SparseEmbedding object as query input, typically from a sparse text embedder. It outputs a list of retrieved Document objects you can connect to PromptBuilder, a ranker, or DocumentJoiner for hybrid retrieval.
Source Code
To check this component's source code, open sparse_embedding_retriever.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
ElasticsearchSparseEmbeddingRetriever:
type: haystack_integrations.components.retrievers.elasticsearch.sparse_embedding_retriever.ElasticsearchSparseEmbeddingRetriever
init_parameters:
top_k: 10
filter_policy: replace
document_store:
type: haystack_integrations.document_stores.elasticsearch.document_store.ElasticsearchDocumentStore
init_parameters:
index: my_index
sparse_vector_field: sparse_vec
embedding_similarity_function: cosine
Using the Component in a Pipeline
# haystack-pipeline
components:
sparse_embedder:
type: haystack_integrations.components.embedders.fastembed.fastembed_sparse_text_embedder.FastembedSparseTextEmbedder
init_parameters:
model: prithvida/Splade_PP_en_v1
retriever:
type: haystack_integrations.components.retrievers.elasticsearch.sparse_embedding_retriever.ElasticsearchSparseEmbeddingRetriever
init_parameters:
top_k: 10
filter_policy: replace
document_store:
type: haystack_integrations.document_stores.elasticsearch.document_store.ElasticsearchDocumentStore
init_parameters:
hosts:
index: my_index
sparse_vector_field: sparse_vec
embedding_similarity_function: cosine
connections:
- sender: sparse_embedder.sparse_embedding
receiver: retriever.query_sparse_embedding
max_runs_per_component: 100
metadata: {}
inputs:
query:
- sparse_embedder.text
outputs:
documents: retriever.documents
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
query_sparse_embedding | SparseEmbedding | The sparse embedding of the query, produced by a sparse text embedder. |
filters | Optional[Dict[str, Any]] | Filters applied when fetching documents. How runtime filters interact with init-time filters depends on filter_policy. |
top_k | Optional[int] | Maximum number of documents to return. |
Outputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | The retrieved documents, ranked by sparse vector similarity. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
document_store | ElasticsearchDocumentStore | An instance of ElasticsearchDocumentStore with sparse_vector_field configured. | |
filters | Optional[Dict[str, Any]] | None | Default filters applied when running the retriever. |
top_k | int | 10 | Maximum number of documents to return. |
filter_policy | Union[str, FilterPolicy] | FilterPolicy.REPLACE | Policy for applying runtime filters. REPLACE overrides init-time filters; MERGE combines them. |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Default | Description |
|---|---|---|---|
query_sparse_embedding | SparseEmbedding | The sparse embedding of the query. | |
filters | Optional[Dict[str, Any]] | None | Runtime filters. Applied according to filter_policy. |
top_k | Optional[int] | None | Override the init-time top_k. |
Related Information
Was this page helpful?