ElasticsearchInferenceHybridRetriever
Combine BM25 keyword search with ELSER sparse vector search in a single server-side request using Elasticsearch's Reciprocal Rank Fusion (RRF)—no local embedding model or client-side score merging required.
Key Features
- Performs hybrid retrieval entirely inside Elasticsearch, using the
retriever.rrfAPI (requires Elasticsearch 8.9+ forrank.rrf, or 8.14+ for the Retriever API). - Uses an Elasticsearch inference endpoint (for example, ELSER) for sparse vector search instead of a client-side embedding model.
- Merges BM25 and sparse results with RRF, with configurable window size and rank constant.
- Supports both synchronous (
run) and asynchronous (run_async) execution. - Configurable filter policy: replace (default) or merge with runtime filters.
Configuration
- Deploy an Elasticsearch inference endpoint in your Elastic cluster (for example, ELSER under the endpoint ID
.elser-2-elasticsearch). - Set your Elasticsearch connection credentials. For instructions, see Create Secrets.
- Configure the
ElasticsearchDocumentStorewithsparse_vector_fieldpointing to the field that stores ELSER sparse vectors. - Drag the
ElasticsearchInferenceHybridRetrievercomponent onto the canvas from the Component Library. - Set
inference_idto the ID of your inference endpoint.
Connections
ElasticsearchInferenceHybridRetriever receives a query string. It outputs a documents list of ranked results. Connect the output to a ranker or directly to the pipeline output. You can pass filters and top_k at run time to override the initialization values.
Source Code
To check this component's source code, open inference_hybrid_retriever.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
ElasticsearchInferenceHybridRetriever:
type: haystack_integrations.components.retrievers.elasticsearch.inference_hybrid_retriever.ElasticsearchInferenceHybridRetriever
init_parameters:
inference_id: .elser-2-elasticsearch
top_k: 10
fuzziness: AUTO
filter_policy: replace
rank_window_size: 100
rank_constant: 60
document_store:
type: haystack_integrations.document_stores.elasticsearch.document_store.ElasticsearchDocumentStore
init_parameters:
hosts:
- https://my-cluster.es.io:9243
api_key:
type: env_var
env_vars:
- ELASTIC_API_KEY
strict: true
index: my-index
sparse_vector_field: sparse_vec
embedding_similarity_function: cosine
Using the Component in a Pipeline
This example shows a server-side hybrid retrieval pipeline that combines keyword and sparse vector search without a local embedding model.
# haystack-pipeline
components:
retriever:
type: haystack_integrations.components.retrievers.elasticsearch.inference_hybrid_retriever.ElasticsearchInferenceHybridRetriever
init_parameters:
inference_id: .elser-2-elasticsearch
top_k: 10
fuzziness: AUTO
rank_window_size: 100
rank_constant: 60
document_store:
type: haystack_integrations.document_stores.elasticsearch.document_store.ElasticsearchDocumentStore
init_parameters:
hosts:
- https://my-cluster.es.io:9243
api_key:
type: env_var
env_vars:
- ELASTIC_API_KEY
strict: true
index: my-index
sparse_vector_field: sparse_vec
embedding_similarity_function: cosine
prompt_builder:
type: haystack.components.builders.chat_prompt_builder.ChatPromptBuilder
init_parameters:
template:
- role: user
content: "Answer the question based on the documents:\n\n{% for doc in documents %}{{ doc.content }}\n\n{% endfor %}\n\nQuestion: {{ query }}"
required_variables:
- documents
- query
llm:
type: haystack.components.generators.chat.openai.OpenAIChatGenerator
init_parameters:
model: gpt-4o-mini
connections:
- sender: retriever.documents
receiver: prompt_builder.documents
- sender: prompt_builder.prompt
receiver: llm.messages
inputs:
query:
- retriever.query
- prompt_builder.query
outputs:
replies: llm.replies
max_runs_per_component: 100
metadata: {}
Parameters
Inputs
| Parameter | Type | Default | Description |
|---|---|---|---|
query | str | The query string to search for. | |
filters | Optional[Dict[str, Any]] | None | Runtime filters applied to both sub-retrievers. How filters are applied depends on filter_policy. |
top_k | Optional[int] | None | Maximum number of documents to return. Overrides the initialization value. |
Outputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | Ranked list of documents combining BM25 and sparse vector results. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
document_store | ElasticsearchDocumentStore | An instance of ElasticsearchDocumentStore with sparse_vector_field configured. | |
inference_id | str | Elasticsearch inference endpoint ID used for sparse vector search (for example, .elser-2-elasticsearch). | |
filters | Optional[Dict[str, Any]] | None | Default filters applied to the retrieved documents. |
fuzziness | str | AUTO | Fuzziness setting for the BM25 keyword query. |
top_k | int | 10 | Maximum number of documents to return. |
filter_policy | Union[str, FilterPolicy] | FilterPolicy.REPLACE | Policy for merging runtime filters with initialization filters. REPLACE overrides init filters; MERGE combines them. |
rank_window_size | int | 100 | Number of candidates each sub-retriever collects before RRF ranking. Higher values improve recall at the cost of latency. |
rank_constant | int | 60 | RRF rank constant. Higher values reduce the impact of rank position differences between sub-retrievers. |
Related Information
Was this page helpful?