Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

TextEmbeddingRetriever

Retrieve documents from a text query using an embedding retriever. The component embeds the query, runs the retriever you configure, and returns documents sorted by score. Use it when a component such as MultiRetriever needs a text query, and your search is embedding-based.

Key Features​

  • Accepts a text query and converts it to an embedding with the text embedder you set.
  • Runs any embedding retriever, such as OpenSearchEmbeddingRetriever.
  • Sorts the returned documents by relevance score.
  • Accepts filters and top_k at query time.

Configuration​

  1. Drag the TextEmbeddingRetriever component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    • Set Retriever to an embedding retriever and choose its document store.
    • Set Text Embedder to the embedder that turns the query into a vector. Use the same model family you used when you embedded the documents.

Connections​

TextEmbeddingRetriever accepts a query string. It outputs documents, a list sorted by score.

Connect query from the pipeline input, or from a component that produces a query string. Connect documents to MultiRetriever, a ranker, or an LLM.

Source Code​

To check this component's source code, open text_embedding_retriever.py in the Haystack repository.

Usage Examples​

Basic Configuration​

TextEmbeddingRetriever:
type: haystack.components.retrievers.text_embedding_retriever.TextEmbeddingRetriever
init_parameters:
retriever:
type: haystack_integrations.components.retrievers.opensearch.embedding_retriever.OpenSearchEmbeddingRetriever
init_parameters:
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
- ${OPENSEARCH_HOST}
index: documents
embedding_dim: 384
top_k: 10
text_embedder:
type: haystack_integrations.components.embedders.sentence_transformers.sentence_transformers_text_embedder.SentenceTransformersTextEmbedder
init_parameters:
model: sentence-transformers/all-MiniLM-L6-v2

Using the Component in a Query Pipeline​

This query pipeline embeds the user query, retrieves matching documents, and sends them to an LLM.

# haystack-pipeline
components:
TextEmbeddingRetriever:
type: haystack.components.retrievers.text_embedding_retriever.TextEmbeddingRetriever
init_parameters:
retriever:
type: haystack_integrations.components.retrievers.opensearch.embedding_retriever.OpenSearchEmbeddingRetriever
init_parameters:
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
- ${OPENSEARCH_HOST}
index: documents
embedding_dim: 384
top_k: 10
text_embedder:
type: haystack_integrations.components.embedders.sentence_transformers.sentence_transformers_text_embedder.SentenceTransformersTextEmbedder
init_parameters:
model: sentence-transformers/all-MiniLM-L6-v2
LLM:
type: haystack.components.generators.chat.llm.LLM
init_parameters:
user_prompt: >-
{% message role="user" %}
Answer the question using the documents.
{% for doc in documents %}
Document {{ loop.index }}: {{ doc.content }}
{% endfor %}
Question: {{ query }}
{% endmessage %}
required_variables: "*"

connections:
- sender: TextEmbeddingRetriever.documents
receiver: LLM.documents

max_runs_per_component: 100

inputs:
query:
- TextEmbeddingRetriever.query
- LLM.query

outputs:
documents: TextEmbeddingRetriever.documents
messages: LLM.messages

metadata: {}

Parameters​

Inputs​

ParameterTypeDescription
querystrThe text query to embed and search with.

Outputs​

ParameterTypeDescription
documentsList[Document]Retrieved documents, sorted by relevance score.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
retrieverEmbeddingRetrieverThe embedding retriever that searches the document store.
text_embedderTextEmbedderThe embedder that converts the query text into a vector.

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDefaultDescription
querystrThe text query to embed and search with.
filtersOptional[Dict]NoneMetadata filters applied by the underlying retriever.
top_kOptional[int]NoneMaximum number of documents to return. When empty, the retriever uses its own top_k.