Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

OpenSearchSQLRetriever

Execute raw SQL queries against an OpenSearchDocumentStore index and return the unprocessed JSON response from the OpenSearch SQL API.

Key Features​

  • Runs arbitrary SQL queries directly against an OpenSearch index using the OpenSearch SQL API.
  • Returns the raw JSON response, giving you access to query hits, aggregations, and other response fields.
  • Configurable page size with fetch_size.
  • Graceful error handling: can either raise on failure or log a warning and return an empty result.
  • Supports both synchronous and asynchronous execution.

Configuration​

  1. Drag the OpenSearchSQLRetriever component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    1. Configure the document_store connection with your OpenSearch instance details.
  4. Go to the Advanced tab to configure fetch_size and raise_on_failure.

Connections​

OpenSearchSQLRetriever receives an SQL query string as input and outputs the raw JSON response from OpenSearch. Connect it to downstream components that consume structured data, such as a custom pipeline output or a data-processing component.

Source Code​

To check this component's source code, open sql_retriever.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

OpenSearchSQLRetriever:
type: haystack_integrations.components.retrievers.opensearch.sql_retriever.OpenSearchSQLRetriever
init_parameters:
raise_on_failure: true
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
index: my-index
embedding_dim: 768

Using the Component in a Pipeline​

# haystack-pipeline
components:
retriever:
type: haystack_integrations.components.retrievers.opensearch.sql_retriever.OpenSearchSQLRetriever
init_parameters:
raise_on_failure: true
fetch_size: 100
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
index: my-index
embedding_dim: 768

connections: []

max_runs_per_component: 100

metadata: {}

inputs:
query:
- retriever.query

outputs:
result: retriever.result

Parameters​

Inputs​

ParameterTypeDescription
querystrThe OpenSearch SQL query to execute, for example SELECT content, category FROM my_index WHERE category = 'A'.
fetch_sizeOptional[int]Number of results to fetch per page. Overrides the init-time value.

Outputs​

ParameterTypeDescription
resultDict[str, Any]The raw JSON response from the OpenSearch SQL API. For regular queries, access documents via result["hits"]["hits"]. For aggregate queries, access results via result["aggregations"].

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
document_storeOpenSearchDocumentStoreAn instance of OpenSearchDocumentStore.
raise_on_failureboolTrueWhen True, raises an exception if the SQL query fails. When False, logs a warning and returns an empty dictionary.
fetch_sizeOptional[int]NoneNumber of results per page. When not set, OpenSearch uses its default page size.

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDefaultDescription
querystrThe OpenSearch SQL query to execute.
fetch_sizeOptional[int]NoneOverride the init-time page size for this request.