Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

MariaDBEmbeddingRetriever

Retrieve documents from a MariaDBDocumentStore using vector similarity search. Use this component in query pipelines to find semantically similar documents using MariaDB's native VECTOR support and MHNSW indexing.

Embedding Models in Pipelines and Indexes

The embedding model you use to embed documents in your index must be the same as the embedding model you use to embed the query in your pipeline.

This means the embedders for your indexes and pipelines must match. For example, if you use CohereDocumentEmbedder to embed your documents, you should use CohereTextEmbedder with the same model to embed your queries.

Key Features​

  • Performs approximate nearest-neighbor (ANN) vector search using MariaDB's native VECTOR datatype with MHNSW indexing.
  • Supports cosine and Euclidean distance functions via VEC_DISTANCE_COSINE and VEC_DISTANCE_EUCLIDEAN.
  • Optional score threshold to exclude documents below a minimum similarity score.
  • Configurable filter policy to merge or replace filters at query time.
  • Requires MariaDB 11.7 or later for native VECTOR support.

Configuration​

Add Workspace-Level Integration​

  1. Click your profile icon and choose Settings.
  2. Go to Workspace>Integrations.
  3. Find the provider you want to connect and click Connect next to them.
  4. Enter the API key and any other required details.
  5. Click Connect. You can use this integration in pipelines and indexes in the current workspace.

Add Organization-Level Integration​

  1. Click your profile icon and choose Settings.
  2. Go to Organization>Integrations.
  3. Find the provider you want to connect and click Connect next to them.
  4. Enter the API key and any other required details.
  5. Click Connect. You can use this integration in pipelines and indexes in all workspaces in the current organization.
  1. Set the MARIADB_USER and MARIADB_PASSWORD environment variables with your MariaDB credentials.
  2. First, configure a MariaDBDocumentStore in your pipeline with create_vector_index: true for ANN search.
  3. Drag the MariaDBEmbeddingRetriever component onto the canvas from the Component Library.
  4. Connect an embedder component to provide query_embedding as input.
  5. Connect the retriever output to downstream components such as PromptBuilder.

Connections​

MariaDBEmbeddingRetriever receives a query_embedding (list of floats) from a text embedder such as SentenceTransformersTextEmbedder. It outputs a list of Document objects you can connect to PromptBuilder or other downstream components.

Source Code​

To check this component's source code, open embedding_retriever.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

MariaDBEmbeddingRetriever:
type: haystack_integrations.components.retrievers.mariadb.embedding_retriever.MariaDBEmbeddingRetriever
init_parameters:
document_store: MariaDBDocumentStore
top_k: 5

Using the Component in a Pipeline​

# haystack-pipeline
components:
text_embedder:
type: haystack_integrations.components.embedders.sentence_transformers.SentenceTransformersTextEmbedder
init_parameters:
model: sentence-transformers/all-MiniLM-L6-v2

document_store:
type: haystack_integrations.document_stores.mariadb.document_store.MariaDBDocumentStore
init_parameters:
host: 127.0.0.1
port: 3306
database: haystack
embedding_dimension: 384
distance: cosine
create_vector_index: true

retriever:
type: haystack_integrations.components.retrievers.mariadb.embedding_retriever.MariaDBEmbeddingRetriever
init_parameters:
document_store: document_store
top_k: 5

connections:
- sender: text_embedder.embedding
receiver: retriever.query_embedding

inputs:
query:
- text_embedder.text

outputs:
documents: retriever.documents

Parameters​

Inputs​

ParameterTypeDescription
query_embeddingList[float]The query embedding vector to search for similar documents.
filtersOptional[Dict[str, Any]]Filters to apply when retrieving documents.
top_kOptional[int]The maximum number of documents to retrieve. Overrides the init-time value.
score_thresholdOptional[float]Minimum score to include a document. Overrides the init-time value.

Outputs​

ParameterTypeDescription
documentsList[Document]A list of the most similar documents from the document store, ordered by similarity.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
document_storeMariaDBDocumentStoreThe MariaDB document store to retrieve documents from.
filtersOptional[Dict[str, Any]]NoneDefault filters to apply when retrieving documents.
top_kint10The maximum number of documents to retrieve.
score_thresholdOptional[float]NoneMinimum similarity score to include a document. Documents below this score are excluded.
filter_policyFilterPolicyFilterPolicy.REPLACEHow to handle filters passed at query time. REPLACE replaces init-time filters; MERGE combines them.

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDefaultDescription
query_embeddingList[float]The embedding vector to search with.
filtersOptional[Dict[str, Any]]NoneRuntime filters to apply.
top_kOptional[int]NoneMaximum number of documents to retrieve. Overrides the init-time value.
score_thresholdOptional[float]NoneMinimum similarity score. Overrides the init-time value.