Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

SentenceTransformersSimilarityRanker

Rank documents based on their semantic similarity to the query using a Sentence Transformers cross-encoder model. Use this component as a reranker after retrieval to improve answer quality.

Key Features​

  • Uses a pre-trained cross-encoder model to score query-document pairs.
  • Configurable number of results returned via top_k.
  • Supports score scaling via Sigmoid activation for normalized similarity scores.
  • Supports filtering by score threshold.
  • Deduplicates documents by ID before ranking.
  • Supports multiple backends: PyTorch, ONNX, and OpenVINO.

Configuration​

  1. Drag the SentenceTransformersSimilarityRanker component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    1. Set the ranking model. Pass a local path or the Hugging Face model name of a cross-encoder model. The default is cross-encoder/ms-marco-MiniLM-L-6-v2.
    2. Set top_k to control how many documents to return.
  4. Go to the Advanced tab to configure scale_score, score_threshold, batch_size, backend, and meta_fields_to_embed.

Connections​

SentenceTransformersSimilarityRanker accepts a query string and a list of documents as inputs. Connect it after a retriever or DocumentJoiner in a query pipeline.

It outputs a ranked list of documents sorted from most to least relevant to the query. Connect its documents output to ChatPromptBuilder, AnswerBuilder, or another downstream component.

Source Code​

To check this component's source code, open sentence_transformers_similarity.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

SentenceTransformersSimilarityRanker:
type: haystack_integrations.components.rankers.sentence_transformers.SentenceTransformersSimilarityRanker
init_parameters:
model: cross-encoder/ms-marco-MiniLM-L-6-v2
top_k: 5
scale_score: true

Using the Component in a Pipeline​

# haystack-pipeline
components:
retriever:
type: haystack.components.retrievers.in_memory.embedding_retriever.InMemoryEmbeddingRetriever
init_parameters:
document_store:
type: haystack.document_stores.in_memory.document_store.InMemoryDocumentStore
init_parameters: {}
top_k: 20

ranker:
type: haystack_integrations.components.rankers.sentence_transformers.SentenceTransformersSimilarityRanker
init_parameters:
model: cross-encoder/ms-marco-MiniLM-L-6-v2
top_k: 5

connections:
- sender: retriever.documents
receiver: ranker.documents

max_runs_per_component: 100

metadata: {}

inputs:
query:
- ranker.query

outputs:
documents: ranker.documents

Parameters​

Inputs​

ParameterTypeDescription
querystrThe input query to compare the documents to.
documentsList[Document]A list of documents to be ranked.
top_kOptional[int]The maximum number of documents to return. Overrides the init-time value.
scale_scoreOptional[bool]If True, scales raw logit predictions using a Sigmoid activation function. Overrides the init-time value.
score_thresholdOptional[float]Return documents only with a score above this threshold. Overrides the init-time value.

Outputs​

ParameterTypeDescription
documentsList[Document]Documents closest to the query, sorted from most similar to least similar.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
modelUnion[str, Path]cross-encoder/ms-marco-MiniLM-L-6-v2The ranking model. Pass a local path or the Hugging Face model name of a cross-encoder model.
deviceOptional[ComponentDevice]NoneThe device on which the model is loaded.
tokenOptional[Secret]Secret.from_env_var(["HF_API_TOKEN", "HF_TOKEN"], strict=False)The API token to download private models from Hugging Face.
top_kint10The maximum number of documents to return per query.
query_prefixstr""A string to add at the beginning of the query text before ranking.
query_suffixstr""A string to add at the end of the query text before ranking.
document_prefixstr""A string to add at the beginning of each document before ranking.
document_suffixstr""A string to add at the end of each document before ranking.
meta_fields_to_embedOptional[List[str]]NoneList of metadata fields to include when ranking each document.
embedding_separatorstr\nSeparator to concatenate metadata fields to the document.
scale_scoreboolTrueIf True, scales raw logit predictions using a Sigmoid activation function.
score_thresholdOptional[float]NoneReturn documents with a score above this threshold only.
trust_remote_codeboolFalseWhether to allow custom models and scripts from Hugging Face.
model_kwargsOptional[Dict[str, Any]]NoneAdditional keyword arguments for the model constructor.
tokenizer_kwargsOptional[Dict[str, Any]]NoneAdditional keyword arguments for the tokenizer.
config_kwargsOptional[Dict[str, Any]]NoneAdditional keyword arguments for the model configuration.
backendLiteral["torch", "onnx", "openvino"]torchThe backend to use for the Sentence Transformers model.
batch_sizeint16The batch size to use for inference.

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDefaultDescription
querystrThe input query to compare the documents to.
documentsList[Document]A list of documents to be ranked.
top_kOptional[int]NoneThe maximum number of documents to return.
scale_scoreOptional[bool]NoneWhether to scale scores with a Sigmoid function.
score_thresholdOptional[float]NoneMinimum score threshold for returned documents.