PyversityRanker
Re-rank documents by balancing relevance and diversity using Pyversity diversification algorithms.
Key Features
- Balances relevance and diversity in ranked document lists using Pyversity's diversification algorithms.
- Supports multiple strategies: DPP (Determinantal Point Process) and MMR (Maximal Marginal Relevance).
- Configurable
diversityparameter to control the trade-off between relevance and diversity on a scale from 0.0 to 1.0. - Documents must have both
scoreandembeddingfields populated, as returned by a dense retriever withreturn_embedding=True. - Accepts per-call overrides of
top_k,strategy, anddiversitywithout changing the component configuration.
Configuration
- Drag the
PyversityRankercomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- On the General tab:
- Set the
strategyto the diversification algorithm:DPP(default) orMMR. - Set the
diversityvalue between0.0(maximum relevance) and1.0(maximum diversity). - Optionally set
top_kto limit the number of returned documents.
- Set the
Connections
PyversityRanker receives a list of Document objects. Each document must have both a score and an embedding populated — documents missing either field are skipped with a warning.
It outputs a re-ranked list of documents, ordered by the chosen diversification algorithm, with updated score values reflecting selection scores.
Connect this ranker after a dense retriever (with return_embedding=True) in a query pipeline.
Source Code
To check this component's source code, open ranker.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
PyversityRanker:
type: haystack_integrations.components.rankers.pyversity.ranker.PyversityRanker
init_parameters:
top_k: 5
strategy: DPP
diversity: 0.5
Using the Component in a Pipeline
This is an example of a query pipeline that retrieves documents using an embedding retriever and then re-ranks them with Pyversity to balance relevance and diversity.
# haystack-pipeline
components:
OpenSearchEmbeddingRetriever:
type: haystack_integrations.components.retrievers.opensearch.embedding_retriever.OpenSearchEmbeddingRetriever
init_parameters:
top_k: 20
filter_policy: replace
document_store:
type: haystack_integrations.document_stores.opensearch.document_store.OpenSearchDocumentStore
init_parameters:
hosts:
index: Standard-Index-English
max_chunk_bytes: 104857600
embedding_dim: 768
return_embedding: true
method:
mappings:
settings:
create_index: true
http_auth:
use_ssl:
verify_certs:
timeout:
TextEmbedder:
type: haystack_integrations.components.embedders.sentence_transformers.SentenceTransformersTextEmbedder
init_parameters:
model: sentence-transformers/all-MiniLM-L6-v2
PyversityRanker:
type: haystack_integrations.components.rankers.pyversity.ranker.PyversityRanker
init_parameters:
top_k: 5
strategy: DPP
diversity: 0.5
connections:
- sender: TextEmbedder.embedding
receiver: OpenSearchEmbeddingRetriever.query_embedding
- sender: OpenSearchEmbeddingRetriever.documents
receiver: PyversityRanker.documents
max_runs_per_component: 100
metadata: {}
inputs:
query:
- TextEmbedder.text
outputs:
documents: PyversityRanker.documents
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | A list of documents to re-rank. Each document must have score and embedding populated. |
top_k | Optional[int] | Overrides the initialized top_k for this call. |
strategy | Optional[Strategy] | Overrides the initialized strategy for this call. |
diversity | Optional[float] | Overrides the initialized diversity for this call. |
Outputs
| Parameter | Type | Description |
|---|---|---|
documents | List[Document] | The re-ranked documents, ordered by the diversification algorithm, with updated score values. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
top_k | Optional[int] | None | Number of documents to return after diversification. If None, all documents are returned in diversified order. |
strategy | Strategy | DPP | Diversification strategy. Supported values are DPP (Determinantal Point Process) and MMR (Maximal Marginal Relevance). |
diversity | float | 0.5 | Trade-off between relevance and diversity in the range [0, 1]. 0.0 keeps only the most relevant documents; 1.0 maximizes diversity regardless of relevance. |
Run Method Parameters
These are the parameters you can configure for the component's run() method. This means you can pass these parameters at query time through the API, in Playground, or when running a job. For details, see Modify Pipeline Parameters at Query Time.
| Parameter | Type | Default | Description |
|---|---|---|---|
documents | List[Document] | A list of documents to re-rank. | |
top_k | Optional[int] | None | Overrides the initialized top_k for this call. |
strategy | Optional[Strategy] | None | Overrides the initialized strategy for this call. |
diversity | Optional[float] | None | Overrides the initialized diversity for this call. |
Related Information
Was this page helpful?