Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

ElasticsearchSQLRetriever

Execute raw SQL queries against an ElasticsearchDocumentStore index and return the unprocessed JSON response from the Elasticsearch SQL API.

Key Features​

  • Runs arbitrary SQL queries directly against an Elasticsearch index using the Elasticsearch SQL API.
  • Returns the raw JSON response with columns metadata and rows data.
  • Configurable page size with fetch_size.
  • Graceful error handling: can either raise on failure or log a warning and return an empty result.
  • Supports both synchronous and asynchronous execution.

Configuration​

  1. Drag the ElasticsearchSQLRetriever component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    1. Configure the document_store connection with your Elasticsearch hosts and index name.
  4. Go to the Advanced tab to configure fetch_size and raise_on_failure.

Connections​

ElasticsearchSQLRetriever receives an SQL query string as input and outputs the raw JSON response from Elasticsearch. Connect it to downstream components that consume structured data, such as a custom pipeline output or a data-processing component.

Source Code​

To check this component's source code, open sql_retriever.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

ElasticsearchSQLRetriever:
type: haystack_integrations.components.retrievers.elasticsearch.sql_retriever.ElasticsearchSQLRetriever
init_parameters:
raise_on_failure: true
document_store:
type: haystack_integrations.document_stores.elasticsearch.document_store.ElasticsearchDocumentStore
init_parameters:
index: my_index
embedding_similarity_function: cosine

Using the Component in a Pipeline​

# haystack-pipeline
components:
retriever:
type: haystack_integrations.components.retrievers.elasticsearch.sql_retriever.ElasticsearchSQLRetriever
init_parameters:
raise_on_failure: true
fetch_size: 100
document_store:
type: haystack_integrations.document_stores.elasticsearch.document_store.ElasticsearchDocumentStore
init_parameters:
hosts:
index: my_index
embedding_similarity_function: cosine

connections: []

max_runs_per_component: 100

metadata: {}

inputs:
query:
- retriever.query

outputs:
result: retriever.result

Parameters​

Inputs​

ParameterTypeDescription
querystrThe Elasticsearch SQL query to execute, for example SELECT content, category FROM "my_index" WHERE category = 'A'.
fetch_sizeOptional[int]Number of results to fetch per page. Overrides the init-time value.

Outputs​

ParameterTypeDescription
resultDict[str, Any]The raw JSON response from the Elasticsearch SQL API. Access column metadata via result["columns"] and data rows via result["rows"].

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
document_storeElasticsearchDocumentStoreAn instance of ElasticsearchDocumentStore.
raise_on_failureboolTrueWhen True, raises an exception if the SQL query fails. When False, logs a warning and returns an empty dictionary.
fetch_sizeOptional[int]NoneNumber of results per page. When not set, Elasticsearch uses its default page size.

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDefaultDescription
querystrThe Elasticsearch SQL query to execute.
fetch_sizeOptional[int]NoneOverride the init-time page size for this request.