Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

ElasticsearchInferenceHybridRetriever

Combine BM25 keyword search with ELSER sparse vector search in a single server-side request using Elasticsearch's Reciprocal Rank Fusion (RRF)—no local embedding model or client-side score merging required.

Key Features​

  • Performs hybrid retrieval entirely inside Elasticsearch, using the retriever.rrf API (requires Elasticsearch 8.9+ for rank.rrf, or 8.14+ for the Retriever API).
  • Uses an Elasticsearch inference endpoint (for example, ELSER) for sparse vector search instead of a client-side embedding model.
  • Merges BM25 and sparse results with RRF, with configurable window size and rank constant.
  • Supports both synchronous (run) and asynchronous (run_async) execution.
  • Configurable filter policy: replace (default) or merge with runtime filters.

Configuration​

  1. Deploy an Elasticsearch inference endpoint in your Elastic cluster (for example, ELSER under the endpoint ID .elser-2-elasticsearch).
  2. Set your Elasticsearch connection credentials. For instructions, see Create Secrets.
  3. Configure the ElasticsearchDocumentStore with sparse_vector_field pointing to the field that stores ELSER sparse vectors.
  4. Drag the ElasticsearchInferenceHybridRetriever component onto the canvas from the Component Library.
  5. Set inference_id to the ID of your inference endpoint.

Connections​

ElasticsearchInferenceHybridRetriever receives a query string. It outputs a documents list of ranked results. Connect the output to a ranker or directly to the pipeline output. You can pass filters and top_k at run time to override the initialization values.

Source Code​

To check this component's source code, open inference_hybrid_retriever.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

ElasticsearchInferenceHybridRetriever:
type: haystack_integrations.components.retrievers.elasticsearch.inference_hybrid_retriever.ElasticsearchInferenceHybridRetriever
init_parameters:
inference_id: .elser-2-elasticsearch
top_k: 10
fuzziness: AUTO
filter_policy: replace
rank_window_size: 100
rank_constant: 60
document_store:
type: haystack_integrations.document_stores.elasticsearch.document_store.ElasticsearchDocumentStore
init_parameters:
hosts:
- https://my-cluster.es.io:9243
api_key:
type: env_var
env_vars:
- ELASTIC_API_KEY
strict: true
index: my-index
sparse_vector_field: sparse_vec
embedding_similarity_function: cosine

Using the Component in a Pipeline​

This example shows a server-side hybrid retrieval pipeline that combines keyword and sparse vector search without a local embedding model.

# haystack-pipeline

components:
retriever:
type: haystack_integrations.components.retrievers.elasticsearch.inference_hybrid_retriever.ElasticsearchInferenceHybridRetriever
init_parameters:
inference_id: .elser-2-elasticsearch
top_k: 10
fuzziness: AUTO
rank_window_size: 100
rank_constant: 60
document_store:
type: haystack_integrations.document_stores.elasticsearch.document_store.ElasticsearchDocumentStore
init_parameters:
hosts:
- https://my-cluster.es.io:9243
api_key:
type: env_var
env_vars:
- ELASTIC_API_KEY
strict: true
index: my-index
sparse_vector_field: sparse_vec
embedding_similarity_function: cosine

prompt_builder:
type: haystack.components.builders.chat_prompt_builder.ChatPromptBuilder
init_parameters:
template:
- role: user
content: "Answer the question based on the documents:\n\n{% for doc in documents %}{{ doc.content }}\n\n{% endfor %}\n\nQuestion: {{ query }}"
required_variables:
- documents
- query

llm:
type: haystack.components.generators.chat.openai.OpenAIChatGenerator
init_parameters:
model: gpt-4o-mini

connections:
- sender: retriever.documents
receiver: prompt_builder.documents
- sender: prompt_builder.prompt
receiver: llm.messages

inputs:
query:
- retriever.query
- prompt_builder.query

outputs:
replies: llm.replies

max_runs_per_component: 100

metadata: {}

Parameters​

Inputs​

ParameterTypeDefaultDescription
querystrThe query string to search for.
filtersOptional[Dict[str, Any]]NoneRuntime filters applied to both sub-retrievers. How filters are applied depends on filter_policy.
top_kOptional[int]NoneMaximum number of documents to return. Overrides the initialization value.

Outputs​

ParameterTypeDescription
documentsList[Document]Ranked list of documents combining BM25 and sparse vector results.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
document_storeElasticsearchDocumentStoreAn instance of ElasticsearchDocumentStore with sparse_vector_field configured.
inference_idstrElasticsearch inference endpoint ID used for sparse vector search (for example, .elser-2-elasticsearch).
filtersOptional[Dict[str, Any]]NoneDefault filters applied to the retrieved documents.
fuzzinessstrAUTOFuzziness setting for the BM25 keyword query.
top_kint10Maximum number of documents to return.
filter_policyUnion[str, FilterPolicy]FilterPolicy.REPLACEPolicy for merging runtime filters with initialization filters. REPLACE overrides init filters; MERGE combines them.
rank_window_sizeint100Number of candidates each sub-retriever collects before RRF ranking. Higher values improve recall at the cost of latency.
rank_constantint60RRF rank constant. Higher values reduce the impact of rank position differences between sub-retrievers.