RagasEvaluator
Evaluate RAG pipeline outputs using Ragas metrics such as faithfulness, context recall, and answer relevancy.
Key Features
- Evaluates RAG pipeline outputs using the Ragas evaluation framework.
- Supports any Ragas metric from
ragas.metrics.collections, including faithfulness, context recall, context precision, answer relevancy, and more. - Accepts multiple metrics in a single component run and returns results for each.
- Supports both synchronous and asynchronous execution with a configurable concurrency limit.
- Each metric must be fully configured at construction time, including the LLM used for scoring.
Configuration
- Drag the
RagasEvaluatorcomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- On the General tab:
- Set
ragas_metricsto a list of Ragas metric instances. Each metric must be configured with its LLM (for example, an OpenAI client).
- Set
- Make sure your OpenAI API key is available via the
OPENAI_API_KEYenvironment variable when using OpenAI-backed metrics.
Connections
RagasEvaluator accepts several optional inputs depending on the metrics you configure:
query: The user's question (required by most metrics).response: A generated response string or a list ofChatMessageobjects.documents: Retrieved documents as a list ofDocumentobjects or strings.reference: A reference answer string.reference_contexts: A list of reference context strings that should have been retrieved.multi_responses: A list of multiple responses generated for the same query.rubrics: A dictionary of evaluation rubrics with score keys and criteria values.
It outputs a result dictionary mapping each metric's name to its MetricResult.
Source Code
To check this component's source code, open evaluator.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
RagasEvaluator:
type: haystack_integrations.components.evaluators.ragas.evaluator.RagasEvaluator
init_parameters:
ragas_metrics:
- type: ragas.metrics.collections.Faithfulness
init_parameters:
llm:
provider: openai
model: gpt-4o-mini
concurrency_limit: 4
Using the Component in a Pipeline
This is an example of an evaluation pipeline that scores generated responses for faithfulness and answer relevancy.
# haystack-pipeline
components:
RagasEvaluator:
type: haystack_integrations.components.evaluators.ragas.evaluator.RagasEvaluator
init_parameters:
ragas_metrics:
- type: ragas.metrics.collections.Faithfulness
init_parameters:
llm:
provider: openai
model: gpt-4o-mini
- type: ragas.metrics.collections.AnswerRelevancy
init_parameters:
llm:
provider: openai
model: gpt-4o-mini
concurrency_limit: 4
connections: []
max_runs_per_component: 100
metadata: {}
inputs:
query:
- RagasEvaluator.query
response:
- RagasEvaluator.response
documents:
- RagasEvaluator.documents
outputs:
result: RagasEvaluator.result
Parameters
Inputs
| Parameter | Type | Description |
|---|---|---|
query | Optional[str] | The input question from the user. |
response | Optional[List[ChatMessage] | str] | The generated response. Accepts either a string or a list of ChatMessage objects. |
documents | Optional[List[Document | str]] | The retrieved documents or context strings. |
reference_contexts | Optional[List[str]] | Reference contexts that should have been retrieved for the query. |
multi_responses | Optional[List[str]] | Multiple responses generated for the query. |
reference | Optional[str] | A reference answer for comparison. |
rubrics | Optional[Dict[str, str]] | Evaluation rubrics where keys represent scores and values represent criteria. |
Outputs
| Parameter | Type | Description |
|---|---|---|
result | Dict[str, MetricResult] | A dictionary mapping each metric's name to its MetricResult. |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
ragas_metrics | List[SimpleBaseMetric] | A list of Ragas metrics from ragas.metrics.collections. Each metric must be fully configured (including its LLM) at construction time. For available metrics, see the Ragas documentation. | |
concurrency_limit | int | 4 | The maximum number of metric evaluations to run concurrently. Used only in the async run_async method. |
Related Information
Was this page helpful?