Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

SolrEmbeddingRetriever

Retrieve documents from a SolrDocumentStore using Solr's {!knn} dense vector search. Use this component in query pipelines to find semantically similar documents based on dense embeddings.

Embedding Models in Pipelines and Indexes

The embedding model you use to embed documents in your index must be the same as the embedding model you use to embed the query in your pipeline.

This means the embedders for your indexes and pipelines must match. For example, if you use CohereDocumentEmbedder to embed your documents, you should use CohereTextEmbedder with the same model to embed your queries.

Key Features​

  • Performs dense vector search using Solr's {!knn} query parser.
  • Filters act as a k-NN graph pre-filter, so the search still returns up to top_k documents even when filters are applied.
  • Requires Solr 9.6 or newer for dependable k-NN pre-filtering support.
  • Configurable filter policy to merge or replace filters at query time.
  • Supports asynchronous retrieval with run_async().

Configuration​

  1. First, configure a SolrDocumentStore with an embedding_dim that matches your embedder model.
  2. Drag the SolrEmbeddingRetriever component onto the canvas from the Component Library.
  3. Connect a text embedder component to provide query_embedding as input.
  4. Set auth credentials as secrets called SOLR_USERNAME and SOLR_PASSWORD. For instructions, see Add Secrets.

Connections​

SolrEmbeddingRetriever receives a query_embedding (list of floats) from a text embedder such as SentenceTransformersTextEmbedder. It outputs a list of Document objects you can connect to PromptBuilder or other downstream components.

Source Code​

To check this component's source code, open embedding_retriever.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

SolrEmbeddingRetriever:
type: haystack_integrations.components.retrievers.solr.embedding_retriever.SolrEmbeddingRetriever
init_parameters:
document_store:
type: haystack_integrations.document_stores.solr.document_store.SolrDocumentStore
init_parameters:
url: http://localhost:8983/solr
core: haystack
embedding_dim: 384
top_k: 10

Using the Component in a Pipeline​

# haystack-pipeline
components:
text_embedder:
type: haystack_integrations.components.embedders.sentence_transformers.SentenceTransformersTextEmbedder
init_parameters:
model: sentence-transformers/all-MiniLM-L6-v2

document_store:
type: haystack_integrations.document_stores.solr.document_store.SolrDocumentStore
init_parameters:
url: http://localhost:8983/solr
core: haystack
embedding_dim: 384
auth:
- type: env_var
env_vars:
- SOLR_USERNAME
strict: false
- type: env_var
env_vars:
- SOLR_PASSWORD
strict: false

retriever:
type: haystack_integrations.components.retrievers.solr.embedding_retriever.SolrEmbeddingRetriever
init_parameters:
document_store: document_store
top_k: 10

connections:
- sender: text_embedder.embedding
receiver: retriever.query_embedding

inputs:
query:
- text_embedder.text

outputs:
documents: retriever.documents

Parameters​

Inputs​

ParameterTypeDescription
query_embeddingList[float]The query embedding vector to search for similar documents.
filtersOptional[Dict[str, Any]]Filters to apply at query time.
top_kOptional[int]Maximum number of documents to retrieve. Overrides the init-time value.

Outputs​

ParameterTypeDescription
documentsList[Document]A list of the most similar documents from the document store.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
document_storeSolrDocumentStoreThe Solr document store to search.
filtersOptional[Dict[str, Any]]NoneDefault filters applied to all searches.
top_kint10Maximum number of documents to return.
filter_policyFilterPolicyFilterPolicy.REPLACEHow to handle filters passed at query time. REPLACE replaces init-time filters; MERGE combines them.
raise_on_failureboolTrueWhether a failing search raises an exception or logs and returns no documents.

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDefaultDescription
query_embeddingList[float]The embedding vector to search with.
filtersOptional[Dict[str, Any]]NoneRuntime filters to apply.
top_kOptional[int]NoneMaximum number of documents to retrieve. Overrides the init-time value.