VespaKeywordRetriever
Retrieve documents from a VespaDocumentStore using lexical keyword search. Use this component in query pipelines that rely on BM25 text matching rather than dense vector similarity.
Key Features
- Performs lexical search using Vespa's BM25 YQL queries.
- Configurable Vespa rank profile for scoring (for example
bm25usingbm25(content)). - Supports Haystack metadata filters translated into Vespa YQL.
- Works with your own Vespa application schema and rank profiles.
- Designed for keyword-heavy retrieval workloads where exact term matching matters.
Configuration
Add Workspace-Level Integration
- Click your profile icon and choose Settings.
- Go to Workspace>Integrations.
- Find the provider you want to connect and click Connect next to them.
- Enter the API key and any other required details.
- Click Connect. You can use this integration in pipelines and indexes in the current workspace.
Add Organization-Level Integration
- Click your profile icon and choose Settings.
- Go to Organization>Integrations.
- Find the provider you want to connect and click Connect next to them.
- Enter the API key and any other required details.
- Click Connect. You can use this integration in pipelines and indexes in all workspaces in the current organization.
- Deploy a Vespa application with a schema that includes a
textfield and a BM25 rank profile. - Configure a
VespaDocumentStorein your pipeline, specifying the Vespaurl,schema, andnamespacematching your application. - Drag the
VespaKeywordRetrievercomponent onto the canvas from the Component Library. - Connect the retriever output to downstream components such as
PromptBuilder.
Connections
VespaKeywordRetriever receives a query string as input. It outputs a list of Document objects ranked by BM25 relevance that you can connect to PromptBuilder or other downstream components.
Source Code
To check this component's source code, open keyword_retriever.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
VespaKeywordRetriever:
type: haystack_integrations.components.retrievers.vespa.keyword_retriever.VespaKeywordRetriever
init_parameters:
document_store: VespaDocumentStore
top_k: 5
ranking: bm25
Using the Component in a Pipeline
# haystack-pipeline
components:
document_store:
type: haystack_integrations.document_stores.vespa.document_store.VespaDocumentStore
init_parameters:
url: http://localhost:8080
schema: doc
namespace: doc
retriever:
type: haystack_integrations.components.retrievers.vespa.keyword_retriever.VespaKeywordRetriever
init_parameters:
document_store: document_store
top_k: 5
ranking: bm25
inputs:
query:
- retriever.query
outputs:
documents: retriever.documents
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
query | str | The keyword query string to search for. |
filters | Optional[Dict[str, Any]] | Filters to apply when retrieving documents. |
top_k | Optional[int] | The maximum number of documents to retrieve. Overrides the init-time value. |
Outputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | A list of documents ranked by BM25 relevance from the document store. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
document_store | VespaDocumentStore | The Vespa document store to retrieve documents from. | |
filters | Optional[Dict[str, Any]] | None | Default filters to apply when retrieving documents. |
top_k | int | 10 | The maximum number of documents to retrieve. |
ranking | Optional[str] | "bm25" | Vespa rank profile for lexical matches. Pass None to use the schema default. |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Default | Description |
|---|---|---|---|
query | str | The keyword query string. | |
filters | Optional[Dict[str, Any]] | None | Runtime filters to apply. |
top_k | Optional[int] | None | Maximum number of documents to retrieve. Overrides the init-time value. |
Related Information
Was this page helpful?