Skip to main content
For the complete documentation index for agents and LLMs, see llms.txt.

DeepEvalEvaluator

Evaluate RAG pipeline outputs using DeepEval metrics such as faithfulness, answer relevancy, and contextual precision.

Key Features​

  • Evaluates RAG pipeline outputs using the DeepEval framework.
  • Supports five metrics: answer_relevancy, faithfulness, contextual_precision, contextual_recall, and contextual_relevance.
  • Returns a score and an optional explanation for each evaluation result.
  • Accepts the metric name as a string or as a DeepEvalMetric enum value.
  • Configurable metric parameters (for example, the LLM model used for evaluation).

Configuration​

  1. Drag the DeepEvalEvaluator component onto the canvas from the Component Library.
  2. Click on the component to open the configuration panel.
  3. On the General tab:
    1. Set metric to the evaluation metric you want to use (for example, faithfulness).
    2. Set metric_params to configure the metric, including the model used for scoring (for example, {"model": "gpt-4o"}).
  4. Make sure your OpenAI API key is available via the OPENAI_API_KEY environment variable, as DeepEval uses an LLM to score responses.

Connections​

DeepEvalEvaluator inputs depend on the metric you choose:

  • answer_relevancy, faithfulness, contextual_relevance: requires questions (List[str]), contexts (List[List[str]]), and responses (List[str]).
  • contextual_precision, contextual_recall: requires the above plus ground_truths (List[str]).

It outputs a nested list of results under the results key. Each result contains name, score, and an optional explanation.

Source Code​

To check this component's source code, open evaluator.py in the Haystack Core Integrations repository.

Usage Examples​

Basic Configuration​

DeepEvalEvaluator:
type: haystack_integrations.components.evaluators.deepeval.evaluator.DeepEvalEvaluator
init_parameters:
metric: faithfulness
metric_params:
model: gpt-4o

Using the Component in a Pipeline​

This is an example of an evaluation pipeline that scores generated answers for faithfulness against the retrieved context.

# haystack-pipeline
components:
DeepEvalEvaluator:
type: haystack_integrations.components.evaluators.deepeval.evaluator.DeepEvalEvaluator
init_parameters:
metric: faithfulness
metric_params:
model: gpt-4o

connections: []

max_runs_per_component: 100

metadata: {}

inputs:
questions:
- DeepEvalEvaluator.questions
contexts:
- DeepEvalEvaluator.contexts
responses:
- DeepEvalEvaluator.responses

outputs:
results: DeepEvalEvaluator.results

Parameters​

Inputs​

The inputs depend on the metric selected. Common inputs for most metrics are listed below.

ParameterTypeDescription
questionsList[str]A list of questions asked in the RAG pipeline.
contextsList[List[str]]A list of context lists, one per question, containing the retrieved passages.
responsesList[str]A list of generated responses, one per question.
ground_truthsList[str]A list of expected answers. Required for contextual_precision and contextual_recall.

Outputs​

ParameterTypeDescription
resultsList[List[Dict[str, Any]]]A nested list of metric results. Each inner list corresponds to one input sample and contains dictionaries with name (metric name), score (float), and explanation (optional string).

Init Parameters​

These are the parameters you can configure in Pipeline Builder:

ParameterTypeDefaultDescription
metricstr | DeepEvalMetricThe metric to use for evaluation. Supported values: answer_relevancy, faithfulness, contextual_precision, contextual_recall, contextual_relevance.
metric_paramsOptional[Dict[str, Any]]NoneParameters passed to the metric constructor. Most metrics require a model key specifying the LLM to use for scoring (for example, {"model": "gpt-4o"}).