OpenSearchSQLRetriever
Execute raw SQL queries against an OpenSearchDocumentStore index and return the unprocessed JSON response from the OpenSearch SQL API.
Key Features
- Runs arbitrary SQL queries directly against an OpenSearch index using the OpenSearch SQL API.
- Returns the raw JSON response, giving you access to query hits, aggregations, and other response fields.
- Configurable page size with
fetch_size. - Graceful error handling: can either raise on failure or log a warning and return an empty result.
- Supports both synchronous and asynchronous execution.
Configuration
- Drag the
OpenSearchSQLRetrievercomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- On the General tab:
- Configure the
document_storeconnection with your OpenSearch instance details.
- Configure the
- Go to the Advanced tab to configure
fetch_sizeandraise_on_failure.
Connections
OpenSearchSQLRetriever receives an SQL query string as input and outputs the raw JSON response from OpenSearch. Connect it to downstream components that consume structured data, such as a custom pipeline output or a data-processing component.
Source Code
To check this component's source code, open sql_retriever.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
OpenSearchSQLRetriever:
type: haystack_integrations.components.retrievers.opensearch.sql_retriever.OpenSearchSQLRetriever
init_parameters:
raise_on_failure: true
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
index: my-index
embedding_dim: 768
Using the Component in a Pipeline
# haystack-pipeline
components:
retriever:
type: haystack_integrations.components.retrievers.opensearch.sql_retriever.OpenSearchSQLRetriever
init_parameters:
raise_on_failure: true
fetch_size: 100
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
index: my-index
embedding_dim: 768
connections: []
max_runs_per_component: 100
metadata: {}
inputs:
query:
- retriever.query
outputs:
result: retriever.result
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
query | str | The OpenSearch SQL query to execute, for example SELECT content, category FROM my_index WHERE category = 'A'. |
fetch_size | Optional[int] | Number of results to fetch per page. Overrides the init-time value. |
Outputs
| Parameter | Type | Description |
|---|---|---|
result | Dict[str, Any] | The raw JSON response from the OpenSearch SQL API. For regular queries, access documents via result["hits"]["hits"]. For aggregate queries, access results via result["aggregations"]. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
document_store | OpenSearchDocumentStore | An instance of OpenSearchDocumentStore. | |
raise_on_failure | bool | True | When True, raises an exception if the SQL query fails. When False, logs a warning and returns an empty dictionary. |
fetch_size | Optional[int] | None | Number of results per page. When not set, OpenSearch uses its default page size. |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Default | Description |
|---|---|---|---|
query | str | The OpenSearch SQL query to execute. | |
fetch_size | Optional[int] | None | Override the init-time page size for this request. |
Related Information
Was this page helpful?