ElasticsearchSQLRetriever
Execute raw SQL queries against an ElasticsearchDocumentStore index and return the unprocessed JSON response from the Elasticsearch SQL API.
Key Features
- Runs arbitrary SQL queries directly against an Elasticsearch index using the Elasticsearch SQL API.
- Returns the raw JSON response with
columnsmetadata androwsdata. - Configurable page size with
fetch_size. - Graceful error handling: can either raise on failure or log a warning and return an empty result.
- Supports both synchronous and asynchronous execution.
Configuration
- Drag the
ElasticsearchSQLRetrievercomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- On the General tab:
- Configure the
document_storeconnection with your Elasticsearch hosts and index name.
- Configure the
- Go to the Advanced tab to configure
fetch_sizeandraise_on_failure.
Connections
ElasticsearchSQLRetriever receives an SQL query string as input and outputs the raw JSON response from Elasticsearch. Connect it to downstream components that consume structured data, such as a custom pipeline output or a data-processing component.
Source Code
To check this component's source code, open sql_retriever.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
ElasticsearchSQLRetriever:
type: haystack_integrations.components.retrievers.elasticsearch.sql_retriever.ElasticsearchSQLRetriever
init_parameters:
raise_on_failure: true
document_store:
type: haystack_integrations.document_stores.elasticsearch.document_store.ElasticsearchDocumentStore
init_parameters:
index: my_index
embedding_similarity_function: cosine
Using the Component in a Pipeline
# haystack-pipeline
components:
retriever:
type: haystack_integrations.components.retrievers.elasticsearch.sql_retriever.ElasticsearchSQLRetriever
init_parameters:
raise_on_failure: true
fetch_size: 100
document_store:
type: haystack_integrations.document_stores.elasticsearch.document_store.ElasticsearchDocumentStore
init_parameters:
hosts:
index: my_index
embedding_similarity_function: cosine
connections: []
max_runs_per_component: 100
metadata: {}
inputs:
query:
- retriever.query
outputs:
result: retriever.result
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
query | str | The Elasticsearch SQL query to execute, for example SELECT content, category FROM "my_index" WHERE category = 'A'. |
fetch_size | Optional[int] | Number of results to fetch per page. Overrides the init-time value. |
Outputs
| Parameter | Type | Description |
|---|---|---|
result | Dict[str, Any] | The raw JSON response from the Elasticsearch SQL API. Access column metadata via result["columns"] and data rows via result["rows"]. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
document_store | ElasticsearchDocumentStore | An instance of ElasticsearchDocumentStore. | |
raise_on_failure | bool | True | When True, raises an exception if the SQL query fails. When False, logs a warning and returns an empty dictionary. |
fetch_size | Optional[int] | None | Number of results per page. When not set, Elasticsearch uses its default page size. |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Default | Description |
|---|---|---|---|
query | str | The Elasticsearch SQL query to execute. | |
fetch_size | Optional[int] | None | Override the init-time page size for this request. |
Related Information
Was this page helpful?