FastembedLateInteractionRanker
Re-rank documents using ColBERT late interaction models via Fastembed for token-level similarity scoring.
Key Features
- Uses ColBERT late interaction (MaxSim) scoring to compute token-level similarity between query and document token embeddings.
- Produces more precise re-ranking than single-vector models, especially for longer documents.
- Runs locally using CPU-optimized ONNX models through Fastembed — no GPU required by default.
- Supports GPU acceleration via ONNX Runtime execution providers such as
CUDAExecutionProvider. - Configurable score threshold to filter out low-scoring documents.
- Supports embedding document metadata fields alongside text for richer ranking context.
Configuration
- Drag the
FastembedLateInteractionRankercomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- On the General tab:
- Set the
model_nameto the ColBERT model you want to use (for example,colbert-ir/colbertv2.0). For a list of supported models, see the Fastembed documentation. - Set
top_kto the number of top documents to return after re-ranking.
- Set the
- Go to the Advanced tab to configure
batch_size,score_threshold,meta_fields_to_embed, andmodel_kwargs.
Connections
FastembedLateInteractionRanker receives a query string and a list of Document objects. It returns the same documents re-ranked by their ColBERT MaxSim score, sorted from most to least relevant, up to top_k results.
Use this component in a query pipeline after a retriever to improve result quality. Connect the retriever's documents output and the query to this ranker, then connect its documents output to a prompt builder or answer builder.
Source Code
To check this component's source code, open late_interaction_ranker.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
FastembedLateInteractionRanker:
type: haystack_integrations.components.rankers.fastembed.late_interaction_ranker.FastembedLateInteractionRanker
init_parameters:
model_name: colbert-ir/colbertv2.0
top_k: 10
batch_size: 64
Using the Component in a Pipeline
This is an example of a query pipeline that retrieves documents using BM25 and then re-ranks them with ColBERT for improved precision.
# haystack-pipeline
components:
BM25Retriever:
type: haystack.components.retrievers.in_memory.bm25_retriever.InMemoryBM25Retriever
init_parameters:
top_k: 20
document_store:
type: haystack.document_stores.in_memory.document_store.InMemoryDocumentStore
init_parameters:
bm25_tokenization_regex: (?u)\b\w\w+\b
bm25_algorithm: BM25L
bm25_parameters:
embedding_similarity_function: dot_product
FastembedLateInteractionRanker:
type: haystack_integrations.components.rankers.fastembed.late_interaction_ranker.FastembedLateInteractionRanker
init_parameters:
model_name: colbert-ir/colbertv2.0
top_k: 5
batch_size: 64
parallel:
local_files_only: false
meta_fields_to_embed:
meta_data_separator: "\n"
score_threshold:
model_kwargs:
connections:
- sender: BM25Retriever.documents
receiver: FastembedLateInteractionRanker.documents
max_runs_per_component: 100
metadata: {}
inputs:
query:
- BM25Retriever.query
- FastembedLateInteractionRanker.query
outputs:
documents: FastembedLateInteractionRanker.documents
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
query | str | The input query to compare documents against. |
documents | List[Document] | A list of documents to re-rank. |
top_k | Optional[int] | Overrides the initialized top_k for this call. |
Outputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | The top-ranked documents sorted from most to least relevant, with updated score values. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
model_name | str | colbert-ir/colbertv2.0 | The Fastembed ColBERT model to use. For a list of supported models, see the Fastembed documentation. |
top_k | int | 10 | The maximum number of documents to return after re-ranking. |
cache_dir | Optional[str] | None | The path to the directory where models are cached. Can also be set using the FASTEMBED_CACHE_PATH environment variable. |
threads | Optional[int] | None | Number of threads for a single ONNX Runtime session. |
batch_size | int | 64 | Number of strings to encode at once. |
parallel | Optional[int] | None | If greater than one, uses data-parallel encoding. Set to 0 to use all available cores. Set to None to use default ONNX Runtime threading. |
local_files_only | bool | False | If True, only uses model files from cache_dir. |
meta_fields_to_embed | Optional[List[str]] | None | List of document metadata fields to concatenate with the document content before ranking. |
meta_data_separator | str | \n | Separator used to concatenate metadata fields with document content. |
score_threshold | Optional[float] | None | If set, only documents with a score above this threshold are returned. ColBERT scores are unnormalized sums and typically range from 3 to 25. |
model_kwargs | Optional[Dict[str, Any]] | None | Additional keyword arguments passed to the Fastembed model. Use {"providers": ["CUDAExecutionProvider"]} to run on an NVIDIA GPU. |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Default | Description |
|---|---|---|---|
query | str | The input query to compare documents against. | |
documents | List[Document] | A list of documents to re-rank. | |
top_k | Optional[int] | None | Overrides the initialized top_k for this call. |
Related Information
Was this page helpful?