Skip to main content
Demonstrates that eval model metrics are accumulated back into the original agent’s run_output when AgentAsJudgeEval is used as a post_hook.
agent_as_judge_eval_metrics.py

Run the Example

1

Set up your virtual environment

2

Install dependencies

3

Export your OpenAI API key

4

Run the example

Save the code above as agent_as_judge_eval_metrics.py, then run:
Full source: cookbook/09_evals/agent_as_judge/agent_as_judge_eval_metrics.py