Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

OpenSearchMetadataRetriever

Search and rank metadata fields stored in an OpenSearchDocumentStore using fuzzy or strict text matching, and return the top matching metadata values.

Key Features​

  • Searches across specified metadata fields using prefix, wildcard, or fuzzy matching.
  • Scores results with Jaccard n-gram similarity computed server-side via an OpenSearch Painless script.
  • Supports both strict mode (prefix and wildcard) and fuzzy mode (Damerau-Levenshtein distance).
  • Boosts exact matches with a configurable exact_match_weight.
  • Accepts comma-separated query parts so multiple terms are searched simultaneously.
  • Returns only the specified metadata fields — document content and other metadata are excluded.

Configuration​

  1. Drag the OpenSearchMetadataRetriever component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    1. Configure the document_store connection with your OpenSearch instance details.
    2. Set metadata_fields to the list of metadata field names you want to search and return.
    3. Set top_k to the maximum number of results to return.
  4. Go to the Advanced tab to configure mode, fuzziness, exact_match_weight, and other scoring parameters.
note

This component only works with string metadata fields. Numeric, boolean, and list-of-non-string fields are not supported because OpenSearch does not index them as text.

Connections​

OpenSearchMetadataRetriever receives a text query string as input. It outputs a list of metadata dictionaries containing only the fields you specified in metadata_fields. Connect it to downstream components that accept structured metadata, such as a custom PromptBuilder template or a pipeline output.

Source Code​

To check this component's source code, open metadata_retriever.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

OpenSearchMetadataRetriever:
type: haystack_integrations.components.retrievers.opensearch.metadata_retriever.OpenSearchMetadataRetriever
init_parameters:
metadata_fields:
- category
- status
top_k: 20
mode: fuzzy
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
index: my-index
embedding_dim: 768

Using the Component in a Pipeline​

# haystack-pipeline
components:
retriever:
type: haystack_integrations.components.retrievers.opensearch.metadata_retriever.OpenSearchMetadataRetriever
init_parameters:
metadata_fields:
- category
- author
- status
top_k: 20
mode: fuzzy
fuzziness: 2
exact_match_weight: 0.6
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
index: my-index
embedding_dim: 768

connections: []

max_runs_per_component: 100

metadata: {}

inputs:
query:
- retriever.query

outputs:
metadata: retriever.metadata

Parameters​

Inputs​

ParameterTypeDescription
querystrThe search query. Supports comma-separated parts; each part is searched across all specified fields.
metadata_fieldsOptional[List[str]]Override the init-time list of metadata fields to search and return.
top_kOptional[int]Maximum number of results to return.
filtersOptional[Dict[str, Any]]Extra filters to apply to the search.

Outputs​

ParameterTypeDescription
metadataList[Dict[str, Any]]The top-k metadata dictionaries, containing only the fields specified in metadata_fields.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
document_storeOpenSearchDocumentStoreAn instance of OpenSearchDocumentStore.
metadata_fieldsList[str]Metadata field names to search within and return. Must be non-empty. Only string-typed fields are supported.
top_kint20Maximum number of top results to return.
exact_match_weightfloat0.6Score boost applied to exact matches. Applied in both strict and fuzzy modes.
modestr"fuzzy"Search mode. "strict" uses prefix and wildcard matching; "fuzzy" uses fuzzy dis_max queries.
fuzzinessint or "AUTO"2Maximum Damerau-Levenshtein edit distance for fuzzy matching. Only applies in "fuzzy" mode.
prefix_lengthint0Number of leading characters that must match exactly before fuzzy matching applies. Only applies in "fuzzy" mode.
max_expansionsint200Maximum number of term variations the fuzzy query can generate. Only applies in "fuzzy" mode.
tie_breakerfloat0.7Weight for other matching clauses in the dis_max query. Only applies in "fuzzy" mode.
jaccard_nint3N-gram size for server-side Jaccard similarity scoring. Larger values favor longer token matches.
raise_on_failureboolTrueWhen True, raises an exception if the query fails. When False, logs a warning and returns an empty list.

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDefaultDescription
querystrThe search query.
metadata_fieldsOptional[List[str]]NoneOverride the init-time metadata fields.
top_kOptional[int]NoneOverride the init-time top_k.
exact_match_weightOptional[float]NoneOverride the init-time exact match weight.
modeOptional[str]NoneOverride the init-time search mode.
fuzzinessOptional[int or str]NoneOverride the init-time fuzziness.
prefix_lengthOptional[int]NoneOverride the init-time prefix length.
max_expansionsOptional[int]NoneOverride the init-time max expansions.
tie_breakerOptional[float]NoneOverride the init-time tie breaker.
jaccard_nOptional[int]NoneOverride the init-time n-gram size.
filtersOptional[Dict[str, Any]]NoneExtra filters to apply to the search.