Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

PyversityRanker

Re-rank documents by balancing relevance and diversity using Pyversity diversification algorithms.

Key Features​

  • Balances relevance and diversity in ranked document lists using Pyversity's diversification algorithms.
  • Supports multiple strategies: DPP (Determinantal Point Process) and MMR (Maximal Marginal Relevance).
  • Configurable diversity parameter to control the trade-off between relevance and diversity on a scale from 0.0 to 1.0.
  • Documents must have both score and embedding fields populated, as returned by a dense retriever with return_embedding=True.
  • Accepts per-call overrides of top_k, strategy, and diversity without changing the component configuration.

Configuration​

  1. Drag the PyversityRanker component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    1. Set the strategy to the diversification algorithm: DPP (default) or MMR.
    2. Set the diversity value between 0.0 (maximum relevance) and 1.0 (maximum diversity).
    3. Optionally set top_k to limit the number of returned documents.

Connections​

PyversityRanker receives a list of Document objects. Each document must have both a score and an embedding populated — documents missing either field are skipped with a warning.

It outputs a re-ranked list of documents, ordered by the chosen diversification algorithm, with updated score values reflecting selection scores.

Connect this ranker after a dense retriever (with return_embedding=True) in a query pipeline.

Source Code​

To check this component's source code, open ranker.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

PyversityRanker:
type: haystack_integrations.components.rankers.pyversity.ranker.PyversityRanker
init_parameters:
top_k: 5
strategy: DPP
diversity: 0.5

Using the Component in a Pipeline​

This is an example of a query pipeline that retrieves documents using an embedding retriever and then re-ranks them with Pyversity to balance relevance and diversity.

# haystack-pipeline
components:
OpenSearchEmbeddingRetriever:
type: haystack_integrations.components.retrievers.opensearch.embedding_retriever.OpenSearchEmbeddingRetriever
init_parameters:
top_k: 20
filter_policy: replace
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
index: Standard-Index-English
max_chunk_bytes: 104857600
embedding_dim: 768
return_embedding: true
method:
mappings:
settings:
create_index: true
http_auth:
use_ssl:
verify_certs:
timeout:

TextEmbedder:
type: haystack_integrations.components.embedders.sentence_transformers.SentenceTransformersTextEmbedder
init_parameters:
model: sentence-transformers/all-MiniLM-L6-v2

PyversityRanker:
type: haystack_integrations.components.rankers.pyversity.ranker.PyversityRanker
init_parameters:
top_k: 5
strategy: DPP
diversity: 0.5

connections:
- sender: TextEmbedder.embedding
receiver: OpenSearchEmbeddingRetriever.query_embedding
- sender: OpenSearchEmbeddingRetriever.documents
receiver: PyversityRanker.documents

max_runs_per_component: 100

metadata: {}

inputs:
query:
- TextEmbedder.text

outputs:
documents: PyversityRanker.documents

Parameters​

Inputs​

ParameterTypeDescription
documentsList[Document]A list of documents to re-rank. Each document must have score and embedding populated.
top_kOptional[int]Overrides the initialized top_k for this call.
strategyOptional[Strategy]Overrides the initialized strategy for this call.
diversityOptional[float]Overrides the initialized diversity for this call.

Outputs​

ParameterTypeDescription
documentsList[Document]The re-ranked documents, ordered by the diversification algorithm, with updated score values.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
top_kOptional[int]NoneNumber of documents to return after diversification. If None, all documents are returned in diversified order.
strategyStrategyDPPDiversification strategy. Supported values are DPP (Determinantal Point Process) and MMR (Maximal Marginal Relevance).
diversityfloat0.5Trade-off between relevance and diversity in the range [0, 1]. 0.0 keeps only the most relevant documents; 1.0 maximizes diversity regardless of relevance.

Run Method Parameters​

These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.

ParameterTypeDefaultDescription
documentsList[Document]A list of documents to re-rank.
top_kOptional[int]NoneOverrides the initialized top_k for this call.
strategyOptional[Strategy]NoneOverrides the initialized strategy for this call.
diversityOptional[float]NoneOverrides the initialized diversity for this call.