Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

FastembedLateInteractionRanker

Re-rank documents using ColBERT late interaction models via Fastembed for token-level similarity scoring.

Key Features​

  • Uses ColBERT late interaction (MaxSim) scoring to compute token-level similarity between query and document token embeddings.
  • Produces more precise re-ranking than single-vector models, especially for longer documents.
  • Runs locally using CPU-optimized ONNX models through Fastembed — no GPU required by default.
  • Supports GPU acceleration via ONNX Runtime execution providers such as CUDAExecutionProvider.
  • Configurable score threshold to filter out low-scoring documents.
  • Supports embedding document metadata fields alongside text for richer ranking context.

Configuration​

  1. Drag the FastembedLateInteractionRanker component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    1. Set the model_name to the ColBERT model you want to use (for example, colbert-ir/colbertv2.0). For a list of supported models, see the Fastembed documentation.
    2. Set top_k to the number of top documents to return after re-ranking.
  4. Go to the Advanced tab to configure batch_size, score_threshold, meta_fields_to_embed, and model_kwargs.

Connections​

FastembedLateInteractionRanker receives a query string and a list of Document objects. It returns the same documents re-ranked by their ColBERT MaxSim score, sorted from most to least relevant, up to top_k results.

Use this component in a query pipeline after a retriever to improve result quality. Connect the retriever's documents output and the query to this ranker, then connect its documents output to a prompt builder or answer builder.

Source Code​

To check this component's source code, open late_interaction_ranker.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

FastembedLateInteractionRanker:
type: haystack_integrations.components.rankers.fastembed.late_interaction_ranker.FastembedLateInteractionRanker
init_parameters:
model_name: colbert-ir/colbertv2.0
top_k: 10
batch_size: 64

Using the Component in a Pipeline​

This is an example of a query pipeline that retrieves documents using BM25 and then re-ranks them with ColBERT for improved precision.

# haystack-pipeline
components:
BM25Retriever:
type: haystack.components.retrievers.in_memory.bm25_retriever.InMemoryBM25Retriever
init_parameters:
top_k: 20
document_store:
type: haystack.document_stores.in_memory.document_store.InMemoryDocumentStore
init_parameters:
bm25_tokenization_regex: (?u)\b\w\w+\b
bm25_algorithm: BM25L
bm25_parameters:
embedding_similarity_function: dot_product

FastembedLateInteractionRanker:
type: haystack_integrations.components.rankers.fastembed.late_interaction_ranker.FastembedLateInteractionRanker
init_parameters:
model_name: colbert-ir/colbertv2.0
top_k: 5
batch_size: 64
parallel:
local_files_only: false
meta_fields_to_embed:
meta_data_separator: "\n"
score_threshold:
model_kwargs:

connections:
- sender: BM25Retriever.documents
receiver: FastembedLateInteractionRanker.documents

max_runs_per_component: 100

metadata: {}

inputs:
query:
- BM25Retriever.query
- FastembedLateInteractionRanker.query

outputs:
documents: FastembedLateInteractionRanker.documents

Parameters​

Inputs​

ParameterTypeDescription
querystrThe input query to compare documents against.
documentsList[Document]A list of documents to re-rank.
top_kOptional[int]Overrides the initialized top_k for this call.

Outputs​

ParameterTypeDescription
documentsList[Document]The top-ranked documents sorted from most to least relevant, with updated score values.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
model_namestrcolbert-ir/colbertv2.0The Fastembed ColBERT model to use. For a list of supported models, see the Fastembed documentation.
top_kint10The maximum number of documents to return after re-ranking.
cache_dirOptional[str]NoneThe path to the directory where models are cached. Can also be set using the FASTEMBED_CACHE_PATH environment variable.
threadsOptional[int]NoneNumber of threads for a single ONNX Runtime session.
batch_sizeint64Number of strings to encode at once.
parallelOptional[int]NoneIf greater than one, uses data-parallel encoding. Set to 0 to use all available cores. Set to None to use default ONNX Runtime threading.
local_files_onlyboolFalseIf True, only uses model files from cache_dir.
meta_fields_to_embedOptional[List[str]]NoneList of document metadata fields to concatenate with the document content before ranking.
meta_data_separatorstr\nSeparator used to concatenate metadata fields with document content.
score_thresholdOptional[float]NoneIf set, only documents with a score above this threshold are returned. ColBERT scores are unnormalized sums and typically range from 3 to 25.
model_kwargsOptional[Dict[str, Any]]NoneAdditional keyword arguments passed to the Fastembed model. Use {"providers": ["CUDAExecutionProvider"]} to run on an NVIDIA GPU.

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDefaultDescription
querystrThe input query to compare documents against.
documentsList[Document]A list of documents to re-rank.
top_kOptional[int]NoneOverrides the initialized top_k for this call.