Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

VespaKeywordRetriever

Retrieve documents from a VespaDocumentStore using lexical keyword search. Use this component in query pipelines that rely on BM25 text matching rather than dense vector similarity.

Key Features​

  • Performs lexical search using Vespa's BM25 YQL queries.
  • Configurable Vespa rank profile for scoring (for example bm25 using bm25(content)).
  • Supports Haystack metadata filters translated into Vespa YQL.
  • Works with your own Vespa application schema and rank profiles.
  • Designed for keyword-heavy retrieval workloads where exact term matching matters.

Configuration​

Add Workspace-Level Integration​

  1. Click your profile icon and choose Settings.
  2. Go to Workspace>Integrations.
  3. Find the provider you want to connect and click Connect next to them.
  4. Enter the API key and any other required details.
  5. Click Connect. You can use this integration in pipelines and indexes in the current workspace.

Add Organization-Level Integration​

  1. Click your profile icon and choose Settings.
  2. Go to Organization>Integrations.
  3. Find the provider you want to connect and click Connect next to them.
  4. Enter the API key and any other required details.
  5. Click Connect. You can use this integration in pipelines and indexes in all workspaces in the current organization.
  1. Deploy a Vespa application with a schema that includes a text field and a BM25 rank profile.
  2. Configure a VespaDocumentStore in your pipeline, specifying the Vespa url, schema, and namespace matching your application.
  3. Drag the VespaKeywordRetriever component onto the canvas from the Component Library.
  4. Connect the retriever output to downstream components such as PromptBuilder.

Connections​

VespaKeywordRetriever receives a query string as input. It outputs a list of Document objects ranked by BM25 relevance that you can connect to PromptBuilder or other downstream components.

Source Code​

To check this component's source code, open keyword_retriever.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

VespaKeywordRetriever:
type: haystack_integrations.components.retrievers.vespa.keyword_retriever.VespaKeywordRetriever
init_parameters:
document_store: VespaDocumentStore
top_k: 5
ranking: bm25

Using the Component in a Pipeline​

# haystack-pipeline
components:
document_store:
type: haystack_integrations.document_stores.vespa.document_store.VespaDocumentStore
init_parameters:
url: http://localhost:8080
schema: doc
namespace: doc

retriever:
type: haystack_integrations.components.retrievers.vespa.keyword_retriever.VespaKeywordRetriever
init_parameters:
document_store: document_store
top_k: 5
ranking: bm25

inputs:
query:
- retriever.query

outputs:
documents: retriever.documents

Parameters​

Inputs​

ParameterTypeDescription
querystrThe keyword query string to search for.
filtersOptional[Dict[str, Any]]Filters to apply when retrieving documents.
top_kOptional[int]The maximum number of documents to retrieve. Overrides the init-time value.

Outputs​

ParameterTypeDescription
documentsList[Document]A list of documents ranked by BM25 relevance from the document store.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
document_storeVespaDocumentStoreThe Vespa document store to retrieve documents from.
filtersOptional[Dict[str, Any]]NoneDefault filters to apply when retrieving documents.
top_kint10The maximum number of documents to retrieve.
rankingOptional[str]"bm25"Vespa rank profile for lexical matches. Pass None to use the schema default.

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDefaultDescription
querystrThe keyword query string.
filtersOptional[Dict[str, Any]]NoneRuntime filters to apply.
top_kOptional[int]NoneMaximum number of documents to retrieve. Overrides the init-time value.