DeepEvalEvaluator
Evaluate RAG pipeline outputs using DeepEval metrics such as faithfulness, answer relevancy, and contextual precision.
Key Features
- Evaluates RAG pipeline outputs using the DeepEval framework.
- Supports five metrics:
answer_relevancy,faithfulness,contextual_precision,contextual_recall, andcontextual_relevance. - Returns a score and an optional explanation for each evaluation result.
- Accepts the metric name as a string or as a
DeepEvalMetricenum value. - Configurable metric parameters (for example, the LLM model used for evaluation).
Configuration
- Drag the
DeepEvalEvaluatorcomponent onto the canvas from the Component Library. - Click on the component to open the configuration panel.
- On the General tab:
- Set
metricto the evaluation metric you want to use (for example,faithfulness). - Set
metric_paramsto configure the metric, including the model used for scoring (for example,{"model": "gpt-4o"}).
- Set
- Make sure your OpenAI API key is available via the
OPENAI_API_KEYenvironment variable, as DeepEval uses an LLM to score responses.
Connections
DeepEvalEvaluator inputs depend on the metric you choose:
answer_relevancy,faithfulness,contextual_relevance: requiresquestions(List[str]),contexts(List[List[str]]), andresponses(List[str]).contextual_precision,contextual_recall: requires the above plusground_truths(List[str]).
It outputs a nested list of results under the results key. Each result contains name, score, and an optional explanation.
Source Code
To check this component's source code, open evaluator.py in the Haystack Core Integrations repository.
Usage Examples
Basic Configuration
DeepEvalEvaluator:
type: haystack_integrations.components.evaluators.deepeval.evaluator.DeepEvalEvaluator
init_parameters:
metric: faithfulness
metric_params:
model: gpt-4o
Using the Component in a Pipeline
This is an example of an evaluation pipeline that scores generated answers for faithfulness against the retrieved context.
# haystack-pipeline
components:
DeepEvalEvaluator:
type: haystack_integrations.components.evaluators.deepeval.evaluator.DeepEvalEvaluator
init_parameters:
metric: faithfulness
metric_params:
model: gpt-4o
connections: []
max_runs_per_component: 100
metadata: {}
inputs:
questions:
- DeepEvalEvaluator.questions
contexts:
- DeepEvalEvaluator.contexts
responses:
- DeepEvalEvaluator.responses
outputs:
results: DeepEvalEvaluator.results
Parameters
Inputs
The inputs depend on the metric selected. Common inputs for most metrics are listed below.
| Parameter | Type | Description |
|---|---|---|
questions | List[str] | A list of questions asked in the RAG pipeline. |
contexts | List[List[str]] | A list of context lists, one per question, containing the retrieved passages. |
responses | List[str] | A list of generated responses, one per question. |
ground_truths | List[str] | A list of expected answers. Required for contextual_precision and contextual_recall. |
Outputs
| Parameter | Type | Description |
|---|---|---|
results | List[List[Dict[str, Any]]] | A nested list of metric results. Each inner list corresponds to one input sample and contains dictionaries with name (metric name), score (float), and explanation (optional string). |
Init Parameters
These are the parameters you can configure in Pipeline Builder:
| Parameter | Type | Default | Description |
|---|---|---|---|
metric | str | DeepEvalMetric | The metric to use for evaluation. Supported values: answer_relevancy, faithfulness, contextual_precision, contextual_recall, contextual_relevance. | |
metric_params | Optional[Dict[str, Any]] | None | Parameters passed to the metric constructor. Most metrics require a model key specifying the LLM to use for scoring (for example, {"model": "gpt-4o"}). |
Related Information
Was this page helpful?