Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

ElasticsearchSparseEmbeddingRetriever

Retrieve documents from an ElasticsearchDocumentStore using sparse vector similarity search with client-side sparse embeddings such as ELSER.

Key Features​

  • Retrieves documents using sparse vector similarity against a pre-indexed sparse_vector_field.
  • Works with any client-side sparse embedding model, including ELSER.
  • Configurable number of results with top_k.
  • Supports metadata filtering to narrow down the search space.
  • Configurable filter policy (replace or merge) for runtime filters.
  • Supports both synchronous and asynchronous execution.

Configuration​

  1. Drag the ElasticsearchSparseEmbeddingRetriever component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    1. Configure the document_store connection. Make sure the ElasticsearchDocumentStore has sparse_vector_field set to the name of the field where your sparse embeddings are stored.
    2. Set top_k to the maximum number of documents to retrieve.
  4. Go to the Advanced tab to configure filters and filter_policy.
note

This retriever accepts a pre-computed SparseEmbedding object as its query input. Use it together with a sparse text embedder such as ELSER to generate the query embedding before retrieval. For server-side inference-based sparse retrieval, use ElasticsearchInferenceSparseRetriever instead.

Connections​

ElasticsearchSparseEmbeddingRetriever receives a SparseEmbedding object as query input, typically from a sparse text embedder. It outputs a list of retrieved Document objects you can connect to PromptBuilder, a ranker, or DocumentJoiner for hybrid retrieval.

Source Code​

To check this component's source code, open sparse_embedding_retriever.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

ElasticsearchSparseEmbeddingRetriever:
type: haystack_integrations.components.retrievers.elasticsearch.sparse_embedding_retriever.ElasticsearchSparseEmbeddingRetriever
init_parameters:
top_k: 10
filter_policy: replace
document_store:
type: haystack_integrations.document_stores.elasticsearch.document_store.ElasticsearchDocumentStore
init_parameters:
index: my_index
sparse_vector_field: sparse_vec
embedding_similarity_function: cosine

Using the Component in a Pipeline​

# haystack-pipeline
components:
sparse_embedder:
type: haystack_integrations.components.embedders.fastembed.fastembed_sparse_text_embedder.FastembedSparseTextEmbedder
init_parameters:
model: prithvida/Splade_PP_en_v1

retriever:
type: haystack_integrations.components.retrievers.elasticsearch.sparse_embedding_retriever.ElasticsearchSparseEmbeddingRetriever
init_parameters:
top_k: 10
filter_policy: replace
document_store:
type: haystack_integrations.document_stores.elasticsearch.document_store.ElasticsearchDocumentStore
init_parameters:
hosts:
index: my_index
sparse_vector_field: sparse_vec
embedding_similarity_function: cosine

connections:
- sender: sparse_embedder.sparse_embedding
receiver: retriever.query_sparse_embedding

max_runs_per_component: 100

metadata: {}

inputs:
query:
- sparse_embedder.text

outputs:
documents: retriever.documents

Parameters​

Inputs​

ParameterTypeDescription
query_sparse_embeddingSparseEmbeddingThe sparse embedding of the query, produced by a sparse text embedder.
filtersOptional[Dict[str, Any]]Filters applied when fetching documents. How runtime filters interact with init-time filters depends on filter_policy.
top_kOptional[int]Maximum number of documents to return.

Outputs​

ParameterTypeDescription
documentsList[Document]The retrieved documents, ranked by sparse vector similarity.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
document_storeElasticsearchDocumentStoreAn instance of ElasticsearchDocumentStore with sparse_vector_field configured.
filtersOptional[Dict[str, Any]]NoneDefault filters applied when running the retriever.
top_kint10Maximum number of documents to return.
filter_policyUnion[str, FilterPolicy]FilterPolicy.REPLACEPolicy for applying runtime filters. REPLACE overrides init-time filters; MERGE combines them.

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDefaultDescription
query_sparse_embeddingSparseEmbeddingThe sparse embedding of the query.
filtersOptional[Dict[str, Any]]NoneRuntime filters. Applied according to filter_policy.
top_kOptional[int]NoneOverride the init-time top_k.