eval/conftest.yml for
the models and prompts, and the evaluate_agent call for how many times each
scenario runs.
Configure the models
eval/conftest.yml at the project root names two models. simulation.llm
plays the user. evaluation.llm is the judge that scores criteria and quality
metrics.
eval/conftest.yml
provider
and model are required; any other key the provider accepts can be added next
to them.
The three prompts
Three built-in prompt templates drive the simulated user and the judge. Each can be replaced with your own Jinja2 file.Override a prompt
Add the path under the matching section ofeval/conftest.yml. Paths are
relative to the project root and cannot leave it. Leave a key out to use the
built-in template.
eval/conftest.yml
Simulated user
Tells the model to stay in character as the user fromsimulation_context,
keep messages short, move toward the goal one step at a time, and end the
conversation once the goal is met or the agent has made clear it cannot help.
It receives simulation_context and returns the next user message and a
done flag.
Built-in simulated user prompt
Built-in simulated user prompt
Criteria judge
Scores each of yourgoals.criteria against the transcript and a ledger of
the events the agent recorded. It receives criteria_text, transcript, and
event_ledger. The template adapts its event vocabulary to the engine of the
running server; this is the text a Mantle run uses.
Built-in criteria judge prompt
Built-in criteria judge prompt
Metrics judge
Scores the built-in quality metrics from the transcript. It receivestranscript.
Built-in metrics judge prompt
Built-in metrics judge prompt
How a run is judged
Criteria
Yourgoals.criteria sentences. The judge receives the transcript, plus a
ledger of the events the agent recorded: which skill activated and completed,
which memory keys were written, which tools ran and whether they returned an
error. It uses the ledger to catch an agent that says “your order is placed”
without having called the tool that places it. Each criterion is returned with
PASS or FAIL and a rationale. Criteria decide whether the run passes.
Built-in quality metrics
The judge also scores every run on the same fixed set of metrics, whether or not you wrote criteria:
The four scores are averaged into
bot_quality. The judge also writes a short
summary of the conversation.
Metrics are recorded, not enforced. A run can pass with task_completion: FAIL if all of its assertions and criteria passed. If task completion must be
required for a scenario, write it as a criterion.
Run the evaluation
Two MCP tools do the work. Your coding agent calls them for you; you can also ask for either one by name.validate_scenario reads one scenario file and reports what is wrong with
it: YAML that does not parse, a missing name or simulation_context, or an
assertion type or field that is not in the scenario schema. Run it after you
edit a file by hand. It does not check assertion keys against the engine;
evaluate_agent does that when it reads the running server.
evaluate_agent runs the scenarios. For each run it:
- Loads
eval/conftest.ymland the scenario, and rejects any assertion key the running engine does not support. - Opens a fresh
sim-<uuid>conversation, sends/session_start, and appliesinitial_slots. - Runs the simulated conversation against your agent.
- Fetches the tracker and evaluates the assertions.
- Sends the transcript and event ledger to the judge for criteria and metrics.
- Writes the result files and updates the batch summary.
run_count is how many times each scenario is simulated. It defaults to 1
and goes up to 10. The simulator is not deterministic, so a scenario that
passes 3 of 3 says more than one that passes 1 of 1. Results are written as
each run finishes; you do not have to wait for the batch to see the first
failure.
Read the results
Each batch writes toeval/results/<timestamp>/, and earlier batches are kept.
Running one scenario three times gives:
run_N.txt is the report you read. run_N.json is
the same run as data, with eval_passed, criteria, task_completion,
latency, scale_metrics, and the assertion rows when the scenario had any.
Use it when you script over results.
A run report
This is what onerun_N.txt contains and what each part means.
overall_resultis the pass rule applied to this run: every assertion and every criterion passed.- Each criterion has the judge’s rationale. Read it before changing anything; it tells you which turn the judge objected to.
- Quality metrics are recorded and do not affect
overall_result. - Each assertion is listed with the id it checked. A failed assertion carries a short detail saying what was found instead.
- The transcript numbers every message. When the agent greets on session
start, the transcript begins with
agent:lines and no leadinguser:line. - Latency is reported for visibility and does not affect pass or fail.
Time to first token (
ttft) is recorded when the server streamed its reply.
The batch summary
summary.txt lists each scenario with its pass count, how long it took, and
where its run files are, then a total across the batch and the time it took.
When any run failed, a final Scenarios requiring attention list names them
with their pass rate. summary.json holds the same numbers as data.
Use a failure as feedback
A failed criterion means the judge decided the handling was wrong; the rationale says where. A failed assertion names the fact that did not hold. If the agent did the right thing in words you did not expect, loosen the criterion. If the agent did the wrong thing, change the skill, its tool constraints, or its instructions, train again, and run the same scenario again.Inspect a conversation
Every report ends with an Inspector URL. Open it to step through that conversation turn by turn and see the memory values and tracker events behind each reply, the same view you get fromrasa inspect.
Simulated conversations use a sim- prefix on the sender id, which sets them
apart from conversations you held yourself.