OpenSearchMetadataRetriever
Search and rank metadata fields stored in an OpenSearchDocumentStore using fuzzy or strict text matching, and return the top matching metadata values.
Key Features
- Searches across specified metadata fields using prefix, wildcard, or fuzzy matching.
- Scores results with Jaccard n-gram similarity computed server-side via an OpenSearch Painless script.
- Supports both
strictmode (prefix and wildcard) andfuzzymode (Damerau-Levenshtein distance). - Boosts exact matches with a configurable
exact_match_weight. - Accepts comma-separated query parts so multiple terms are searched simultaneously.
- Returns only the specified metadata fields — document content and other metadata are excluded.
Configuration
- Drag the
OpenSearchMetadataRetrievercomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- On the General tab:
- Configure the
document_storeconnection with your OpenSearch instance details. - Set
metadata_fieldsto the list of metadata field names you want to search and return. - Set
top_kto the maximum number of results to return.
- Configure the
- Go to the Advanced tab to configure
mode,fuzziness,exact_match_weight, and other scoring parameters.
This component only works with string metadata fields. Numeric, boolean, and list-of-non-string fields are not supported because OpenSearch does not index them as text.
Connections
OpenSearchMetadataRetriever receives a text query string as input. It outputs a list of metadata dictionaries containing only the fields you specified in metadata_fields. Connect it to downstream components that accept structured metadata, such as a custom PromptBuilder template or a pipeline output.
Source Code
To check this component's source code, open metadata_retriever.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
OpenSearchMetadataRetriever:
type: haystack_integrations.components.retrievers.opensearch.metadata_retriever.OpenSearchMetadataRetriever
init_parameters:
metadata_fields:
- category
- status
top_k: 20
mode: fuzzy
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
index: my-index
embedding_dim: 768
Using the Component in a Pipeline
# haystack-pipeline
components:
retriever:
type: haystack_integrations.components.retrievers.opensearch.metadata_retriever.OpenSearchMetadataRetriever
init_parameters:
metadata_fields:
- category
- author
- status
top_k: 20
mode: fuzzy
fuzziness: 2
exact_match_weight: 0.6
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
index: my-index
embedding_dim: 768
connections: []
max_runs_per_component: 100
metadata: {}
inputs:
query:
- retriever.query
outputs:
metadata: retriever.metadata
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
query | str | The search query. Supports comma-separated parts; each part is searched across all specified fields. |
metadata_fields | Optional[List[str]] | Override the init-time list of metadata fields to search and return. |
top_k | Optional[int] | Maximum number of results to return. |
filters | Optional[Dict[str, Any]] | Extra filters to apply to the search. |
Outputs
| Parameter | Type | Description |
|---|---|---|
metadata | List[Dict[str, Any]] | The top-k metadata dictionaries, containing only the fields specified in metadata_fields. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
document_store | OpenSearchDocumentStore | An instance of OpenSearchDocumentStore. | |
metadata_fields | List[str] | Metadata field names to search within and return. Must be non-empty. Only string-typed fields are supported. | |
top_k | int | 20 | Maximum number of top results to return. |
exact_match_weight | float | 0.6 | Score boost applied to exact matches. Applied in both strict and fuzzy modes. |
mode | str | "fuzzy" | Search mode. "strict" uses prefix and wildcard matching; "fuzzy" uses fuzzy dis_max queries. |
fuzziness | int or "AUTO" | 2 | Maximum Damerau-Levenshtein edit distance for fuzzy matching. Only applies in "fuzzy" mode. |
prefix_length | int | 0 | Number of leading characters that must match exactly before fuzzy matching applies. Only applies in "fuzzy" mode. |
max_expansions | int | 200 | Maximum number of term variations the fuzzy query can generate. Only applies in "fuzzy" mode. |
tie_breaker | float | 0.7 | Weight for other matching clauses in the dis_max query. Only applies in "fuzzy" mode. |
jaccard_n | int | 3 | N-gram size for server-side Jaccard similarity scoring. Larger values favor longer token matches. |
raise_on_failure | bool | True | When True, raises an exception if the query fails. When False, logs a warning and returns an empty list. |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Default | Description |
|---|---|---|---|
query | str | The search query. | |
metadata_fields | Optional[List[str]] | None | Override the init-time metadata fields. |
top_k | Optional[int] | None | Override the init-time top_k. |
exact_match_weight | Optional[float] | None | Override the init-time exact match weight. |
mode | Optional[str] | None | Override the init-time search mode. |
fuzziness | Optional[int or str] | None | Override the init-time fuzziness. |
prefix_length | Optional[int] | None | Override the init-time prefix length. |
max_expansions | Optional[int] | None | Override the init-time max expansions. |
tie_breaker | Optional[float] | None | Override the init-time tie breaker. |
jaccard_n | Optional[int] | None | Override the init-time n-gram size. |
filters | Optional[Dict[str, Any]] | None | Extra filters to apply to the search. |
Related Information
Was this page helpful?