DynamoDBEmbeddingRetriever
Retrieve documents from a DynamoDBDocumentStore using Amazon DynamoDB native vector search. Use this component in query pipelines to find semantically similar documents with cosine similarity on stored embeddings.
The embedding model you use to embed documents in your index must be the same as the embedding model you use to embed the query in your pipeline.
This means the embedders for your indexes and pipelines must match. For example, if you use CohereDocumentEmbedder to embed your documents, you should use CohereTextEmbedder with the same model to embed your queries.
Key Features
- Uses DynamoDB
SearchVectorsfor approximate nearest-neighbor retrieval (up to 100 candidates per request). - Applies Haystack metadata filters client-side after vector search.
- Configurable filter policy to merge or replace filters at query time.
- Requires a DynamoDB table with a vector index and matching embedding dimension.
Configuration
Add Workspace-Level Integration
- Click your profile icon and choose Settings.
- Go to Workspace>Integrations.
- Find the provider you want to connect and click Connect next to them.
- Enter the API key and any other required details.
- Click Connect. You can use this integration in pipelines and indexes in the current workspace.
Add Organization-Level Integration
- Click your profile icon and choose Settings.
- Go to Organization>Integrations.
- Find the provider you want to connect and click Connect next to them.
- Enter the API key and any other required details.
- Click Connect. You can use this integration in pipelines and indexes in all workspaces in the current organization.
- Create a DynamoDB table and vector index with an embedding field that matches your pipeline dimension.
- Configure a
DynamoDBDocumentStorein your pipeline with the table name, index name, andembedding_dimension. - Drag the
DynamoDBEmbeddingRetrievercomponent onto the canvas from the Component Library. - Connect an embedder component to provide
query_embeddingas input. - Connect the retriever output to downstream components such as
PromptBuilder.
Connections
DynamoDBEmbeddingRetriever receives a query_embedding (list of floats) from a text embedder such as SentenceTransformersTextEmbedder. It outputs a list of Document objects you can connect to PromptBuilder or other downstream components.
Source Code
To check this component's source code, open embedding_retriever.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
DynamoDBEmbeddingRetriever:
type: haystack_integrations.components.retrievers.dynamodb.embedding_retriever.DynamoDBEmbeddingRetriever
init_parameters:
document_store: DynamoDBDocumentStore
top_k: 5
Using the Component in a Pipeline
# haystack-pipeline
components:
text_embedder:
type: haystack_integrations.components.embedders.sentence_transformers.SentenceTransformersTextEmbedder
init_parameters:
model: sentence-transformers/all-MiniLM-L6-v2
document_store:
type: haystack_integrations.document_stores.dynamodb.document_store.DynamoDBDocumentStore
init_parameters:
table_name: haystack_documents
index_name: haystack_vector_index
embedding_dimension: 384
region_name: us-east-1
retriever:
type: haystack_integrations.components.retrievers.dynamodb.embedding_retriever.DynamoDBEmbeddingRetriever
init_parameters:
document_store: document_store
top_k: 5
connections:
- sender: text_embedder.embedding
receiver: retriever.query_embedding
inputs:
query:
- text_embedder.text
outputs:
documents: retriever.documents
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
query_embedding | List[float] | The query embedding vector used for similarity search. |
top_k | Optional[int] | Maximum number of documents to return. Overrides the init-time value. Must be between 1 and 100. |
filters | Optional[Dict[str, Any]] | Metadata filters applied after vector search. |
Outputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | Documents ranked by vector similarity. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
document_store | DynamoDBDocumentStore | The DynamoDB document store to retrieve documents from. | |
top_k | int | 10 | Maximum number of documents to return (1 to 100). |
filters | Optional[Dict[str, Any]] | None | Default metadata filters applied at retrieval time. |
filter_policy | str | replace | How run-time filters combine with init-time filters (replace or merge). |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Default | Description |
|---|---|---|---|
query_embedding | List[float] | The query embedding vector. | |
top_k | Optional[int] | None | Maximum number of documents to return. Overrides the init-time value. |
filters | Optional[Dict[str, Any]] | None | Runtime metadata filters. |
Related Information
Was this page helpful?