Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

ElasticsearchInferenceSparseRetriever

Retrieve documents using sparse vector search powered by an Elasticsearch inference endpoint such as ELSER—without a local embedding model.

Key Features​

  • Uses an Elasticsearch inference endpoint to generate sparse vectors at query time entirely on the server side.
  • No local embedding model or client-side vector computation required.
  • Compatible with ELSER and any other sparse inference model deployed in your Elastic cluster.
  • Supports both synchronous (run) and asynchronous (run_async) execution.
  • Configurable filter policy: replace (default) or merge with runtime filters.
  • Can be combined with ElasticsearchBM25Retriever and a DocumentJoiner for hybrid retrieval.

Configuration​

  1. Deploy a sparse inference endpoint in your Elastic cluster (for example, ELSER).
  2. Set your Elasticsearch connection credentials. For instructions, see Create Secrets.
  3. Configure the ElasticsearchDocumentStore with sparse_vector_field pointing to the field that stores the sparse vectors in your index.
  4. Drag the ElasticsearchInferenceSparseRetriever component onto the canvas from the Component Library.
  5. Set inference_id to the ID of your inference endpoint.

Connections​

ElasticsearchInferenceSparseRetriever receives a query string. It outputs a documents list of the most relevant results. Connect the output to a ranker, DocumentJoiner, or directly to the pipeline output. You can pass filters and top_k at run time to override the initialization values.

Source Code​

To check this component's source code, open inference_sparse_retriever.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

ElasticsearchInferenceSparseRetriever:
type: haystack_integrations.components.retrievers.elasticsearch.inference_sparse_retriever.ElasticsearchInferenceSparseRetriever
init_parameters:
inference_id: ELSER
top_k: 10
filter_policy: replace
document_store:
type: haystack_integrations.document_stores.elasticsearch.document_store.ElasticsearchDocumentStore
init_parameters:
hosts:
- https://my-cluster.es.io:9243
api_key:
type: env_var
env_vars:
- ELASTIC_API_KEY
strict: true
index: my-index
sparse_vector_field: sparse_vec
embedding_similarity_function: cosine

Using the Component in a Pipeline​

This example shows a semantic search pipeline using server-side sparse vector retrieval.

# haystack-pipeline

components:
retriever:
type: haystack_integrations.components.retrievers.elasticsearch.inference_sparse_retriever.ElasticsearchInferenceSparseRetriever
init_parameters:
inference_id: ELSER
top_k: 10
document_store:
type: haystack_integrations.document_stores.elasticsearch.document_store.ElasticsearchDocumentStore
init_parameters:
hosts:
- https://my-cluster.es.io:9243
api_key:
type: env_var
env_vars:
- ELASTIC_API_KEY
strict: true
index: my-index
sparse_vector_field: sparse_vec
embedding_similarity_function: cosine

prompt_builder:
type: haystack.components.builders.chat_prompt_builder.ChatPromptBuilder
init_parameters:
template:
- role: user
content: "Answer the question using the documents below:\n\n{% for doc in documents %}{{ doc.content }}\n\n{% endfor %}\n\nQuestion: {{ query }}"
required_variables:
- documents
- query

llm:
type: haystack.components.generators.chat.openai.OpenAIChatGenerator
init_parameters:
model: gpt-4o-mini

connections:
- sender: retriever.documents
receiver: prompt_builder.documents
- sender: prompt_builder.prompt
receiver: llm.messages

inputs:
query:
- retriever.query
- prompt_builder.query

outputs:
replies: llm.replies

max_runs_per_component: 100

metadata: {}

Parameters​

Inputs​

ParameterTypeDefaultDescription
querystrThe query string for sparse vector retrieval.
filtersOptional[Dict[str, Any]]NoneRuntime filters applied to the retrieved documents. How filters are applied depends on filter_policy.
top_kOptional[int]NoneMaximum number of documents to return. Overrides the initialization value.

Outputs​

ParameterTypeDescription
documentsList[Document]List of documents most similar to the query, ranked by sparse vector similarity.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
document_storeElasticsearchDocumentStoreAn instance of ElasticsearchDocumentStore with sparse_vector_field configured.
inference_idstrThe Elasticsearch inference model identifier for sparse vector search (for example, ELSER).
filtersOptional[Dict[str, Any]]NoneDefault filters applied to the retrieved documents.
top_kint10Maximum number of documents to return.
filter_policyUnion[str, FilterPolicy]FilterPolicy.REPLACEPolicy for merging runtime filters with initialization filters. REPLACE overrides init filters; MERGE combines them.