ValkeyEmbeddingRetriever
Retrieve documents from a ValkeyDocumentStore using vector similarity search. Use this component in query pipelines to find semantically similar documents stored in Valkey, an open-source Redis-compatible key-value store with vector search capabilities.
The embedding model you use to embed documents in your index must be the same as the embedding model you use to embed the query in your pipeline.
This means the embedders for your indexes and pipelines must match. For example, if you use CohereDocumentEmbedder to embed your documents, you should use CohereTextEmbedder with the same model to embed your queries.
Key Features
- Performs vector similarity search using Valkey Search's HNSW algorithm.
- Supports L2, cosine, and inner product (
ip) distance metrics. - Supports metadata filtering on tag (string) and numeric fields.
- Supports both standalone and cluster mode connections.
- Supports both synchronous and asynchronous execution.
Configuration
Add Workspace-Level Integration
- Click your profile icon and choose Settings.
- Go to Workspace>Integrations.
- Find the provider you want to connect and click Connect next to them.
- Enter the API key and any other required details.
- Click Connect. You can use this integration in pipelines and indexes in the current workspace.
Add Organization-Level Integration
- Click your profile icon and choose Settings.
- Go to Organization>Integrations.
- Find the provider you want to connect and click Connect next to them.
- Enter the API key and any other required details.
- Click Connect. You can use this integration in pipelines and indexes in all workspaces in the current organization.
- Optionally set
VALKEY_USERNAMEandVALKEY_PASSWORDenvironment variables if your Valkey instance requires authentication. - Configure a
ValkeyDocumentStorein your pipeline, specifyingindex_name,embedding_dim, and anymetadata_fieldsyou want to filter on. - Drag the
ValkeyEmbeddingRetrievercomponent onto the canvas from the Component Library. - Connect an embedder component to provide
query_embeddingas input. - Connect the retriever output to downstream components such as
PromptBuilder.
Connections
ValkeyEmbeddingRetriever receives a query_embedding (list of floats) from a text embedder such as SentenceTransformersTextEmbedder. It outputs a list of Document objects you can connect to PromptBuilder or other downstream components.
Source Code
To check this component's source code, open embedding_retriever.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
ValkeyEmbeddingRetriever:
type: haystack_integrations.components.retrievers.valkey.embedding_retriever.ValkeyEmbeddingRetriever
init_parameters:
document_store: ValkeyDocumentStore
top_k: 5
Using the Component in a Pipeline
# haystack-pipeline
components:
text_embedder:
type: haystack_integrations.components.embedders.sentence_transformers.SentenceTransformersTextEmbedder
init_parameters:
model: sentence-transformers/all-MiniLM-L6-v2
document_store:
type: haystack_integrations.document_stores.valkey.document_store.ValkeyDocumentStore
init_parameters:
nodes_list:
- - localhost
- 6379
index_name: haystack_documents
embedding_dim: 384
distance_metric: cosine
retriever:
type: haystack_integrations.components.retrievers.valkey.embedding_retriever.ValkeyEmbeddingRetriever
init_parameters:
document_store: document_store
top_k: 5
connections:
- sender: text_embedder.embedding
receiver: retriever.query_embedding
inputs:
query:
- text_embedder.text
outputs:
documents: retriever.documents
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
query_embedding | List[float] | The query embedding vector to search for similar documents. |
filters | Optional[Dict[str, Any]] | Filters to apply when retrieving documents. Only fields declared in metadata_fields on the document store are filterable. |
top_k | Optional[int] | The maximum number of documents to retrieve. Overrides the init-time value. |
Outputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | A list of the most similar documents from the document store, ordered by similarity score. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
document_store | ValkeyDocumentStore | The Valkey document store to retrieve documents from. | |
filters | Optional[Dict[str, Any]] | None | Default filters to apply when retrieving documents. |
top_k | int | 10 | The maximum number of documents to retrieve. |
filter_policy | FilterPolicy | FilterPolicy.REPLACE | How to handle filters passed at query time. REPLACE replaces init-time filters; MERGE combines them. |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Default | Description |
|---|---|---|---|
query_embedding | List[float] | The embedding vector to search with. | |
filters | Optional[Dict[str, Any]] | None | Runtime filters to apply. |
top_k | Optional[int] | None | Maximum number of documents to retrieve. Overrides the init-time value. |
Related Information
Was this page helpful?