IBMDb2EmbeddingRetriever
Retrieve documents from an IBMDb2DocumentStore using IBM Db2's native vector search capabilities.
The embedding model you use to embed documents in your index must be the same as the embedding model you use to embed the query in your pipeline.
This means the embedders for your indexes and pipelines must match. For example, if you use CohereDocumentEmbedder to embed your documents, you should use CohereTextEmbedder with the same model to embed your queries.
Key Features
- Performs vector similarity search using IBM Db2's native VECTOR type and distance functions.
- Supports COSINE, EUCLIDEAN, and MANHATTAN distance metrics configured at the document store level.
- Configurable filter policy to merge or replace filters at query time.
- Works in Haystack pipelines after any text embedder component.
Configuration
- First, configure an
IBMDb2DocumentStorewith your Db2 connection details. - Drag the
IBMDb2EmbeddingRetrievercomponent onto the canvas from the Component Library. - Connect a text embedder component to provide
query_embeddingas input. - On the General tab, configure the nested
IBMDb2DocumentStore:- Set
databaseandhostnamefor your Db2 server. - Set credentials as secrets called
DB2_USERNAMEandDB2_PASSWORD. For instructions, see Add Secrets.
- Set
Connections
IBMDb2EmbeddingRetriever receives a query_embedding (list of floats) from a text embedder. It outputs a list of Document objects you can connect to PromptBuilder or other downstream components.
Source Code
To check this component's source code, open embedding_retriever.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
IBMDb2EmbeddingRetriever:
type: haystack_integrations.components.retrievers.ibm_db.embedding_retriever.IBMDb2EmbeddingRetriever
init_parameters:
document_store:
type: haystack_integrations.document_stores.ibm_db.document_store.IBMDb2DocumentStore
init_parameters:
database: mydb
hostname: localhost
username:
type: env_var
env_vars:
- DB2_USERNAME
strict: false
password:
type: env_var
env_vars:
- DB2_PASSWORD
strict: false
embedding_dim: 384
top_k: 10
Using the Component in a Pipeline
# haystack-pipeline
components:
text_embedder:
type: haystack_integrations.components.embedders.sentence_transformers.SentenceTransformersTextEmbedder
init_parameters:
model: sentence-transformers/all-MiniLM-L6-v2
document_store:
type: haystack_integrations.document_stores.ibm_db.document_store.IBMDb2DocumentStore
init_parameters:
database: mydb
hostname: db2.example.com
port: 50000
username:
type: env_var
env_vars:
- DB2_USERNAME
strict: false
password:
type: env_var
env_vars:
- DB2_PASSWORD
strict: false
embedding_dim: 384
distance_metric: COSINE
retriever:
type: haystack_integrations.components.retrievers.ibm_db.embedding_retriever.IBMDb2EmbeddingRetriever
init_parameters:
document_store: document_store
top_k: 5
connections:
- sender: text_embedder.embedding
receiver: retriever.query_embedding
inputs:
query:
- text_embedder.text
outputs:
documents: retriever.documents
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
query_embedding | List[float] | The query embedding vector to search for similar documents. |
filters | Optional[Dict[str, Any]] | Filters to apply at query time. |
top_k | Optional[int] | Maximum number of documents to retrieve. Overrides the init-time value. |
Outputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | A list of the most similar documents from the document store. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
document_store | IBMDb2DocumentStore | The IBM Db2 document store to retrieve documents from. | |
filters | Optional[Dict[str, Any]] | None | Default filters applied to all searches. |
top_k | int | 10 | Maximum number of documents to return. |
filter_policy | FilterPolicy | FilterPolicy.REPLACE | How to handle filters passed at query time. REPLACE replaces init-time filters; MERGE combines them. |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Default | Description |
|---|---|---|---|
query_embedding | List[float] | The embedding vector to search with. | |
filters | Optional[Dict[str, Any]] | None | Runtime filters to apply. |
top_k | Optional[int] | None | Maximum number of documents to retrieve. Overrides the init-time value. |
Related Information
Was this page helpful?