Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

RagasEvaluator

Evaluate RAG pipeline outputs using Ragas metrics such as faithfulness, context recall, and answer relevancy.

Key Features​

  • Evaluates RAG pipeline outputs using the Ragas evaluation framework.
  • Supports any Ragas metric from ragas.metrics.collections, including faithfulness, context recall, context precision, answer relevancy, and more.
  • Accepts multiple metrics in a single component run and returns results for each.
  • Supports both synchronous and asynchronous execution with a configurable concurrency limit.
  • Each metric must be fully configured at construction time, including the LLM used for scoring.

Configuration​

  1. Drag the RagasEvaluator component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    1. Set ragas_metrics to a list of Ragas metric instances. Each metric must be configured with its LLM (for example, an OpenAI client).
  4. Make sure your OpenAI API key is available via the OPENAI_API_KEY environment variable when using OpenAI-backed metrics.

Connections​

RagasEvaluator accepts several optional inputs depending on the metrics you configure:

  • query: The user's question (required by most metrics).
  • response: A generated response string or a list of ChatMessage objects.
  • documents: Retrieved documents as a list of Document objects or strings.
  • reference: A reference answer string.
  • reference_contexts: A list of reference context strings that should have been retrieved.
  • multi_responses: A list of multiple responses generated for the same query.
  • rubrics: A dictionary of evaluation rubrics with score keys and criteria values.

It outputs a result dictionary mapping each metric's name to its MetricResult.

Source Code​

To check this component's source code, open evaluator.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

RagasEvaluator:
type: haystack_integrations.components.evaluators.ragas.evaluator.RagasEvaluator
init_parameters:
ragas_metrics:
- type: ragas.metrics.collections.Faithfulness
init_parameters:
llm:
provider: openai
model: gpt-4o-mini
concurrency_limit: 4

Using the Component in a Pipeline​

This is an example of an evaluation pipeline that scores generated responses for faithfulness and answer relevancy.

# haystack-pipeline
components:
RagasEvaluator:
type: haystack_integrations.components.evaluators.ragas.evaluator.RagasEvaluator
init_parameters:
ragas_metrics:
- type: ragas.metrics.collections.Faithfulness
init_parameters:
llm:
provider: openai
model: gpt-4o-mini
- type: ragas.metrics.collections.AnswerRelevancy
init_parameters:
llm:
provider: openai
model: gpt-4o-mini
concurrency_limit: 4

connections: []

max_runs_per_component: 100

metadata: {}

inputs:
query:
- RagasEvaluator.query
response:
- RagasEvaluator.response
documents:
- RagasEvaluator.documents

outputs:
result: RagasEvaluator.result

Parameters​

Inputs​

ParameterTypeDescription
queryOptional[str]The input question from the user.
responseOptional[List[ChatMessage] | str]The generated response. Accepts either a string or a list of ChatMessage objects.
documentsOptional[List[Document | str]]The retrieved documents or context strings.
reference_contextsOptional[List[str]]Reference contexts that should have been retrieved for the query.
multi_responsesOptional[List[str]]Multiple responses generated for the query.
referenceOptional[str]A reference answer for comparison.
rubricsOptional[Dict[str, str]]Evaluation rubrics where keys represent scores and values represent criteria.

Outputs​

ParameterTypeDescription
resultDict[str, MetricResult]A dictionary mapping each metric's name to its MetricResult.

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
ragas_metricsList[SimpleBaseMetric]A list of Ragas metrics from ragas.metrics.collections. Each metric must be fully configured (including its LLM) at construction time. For available metrics, see the Ragas documentation.
concurrency_limitint4The maximum number of metric evaluations to run concurrently. Used only in the async run_async method.