Skip to main content
Everything you can set for a run lives in two places: eval/conftest.yml for the models and prompts, and the evaluate_agent call for how many times each scenario runs.

Configure the models

eval/conftest.yml at the project root names two models. simulation.llm plays the user. evaluation.llm is the judge that scores criteria and quality metrics.
eval/conftest.yml
Configure them independently. The simulator writes short user messages, so a smaller, cheaper model is enough. The judge has to read a whole transcript and decide whether each criterion holds, so give it a more capable model. provider and model are required; any other key the provider accepts can be added next to them.

The three prompts

Three built-in prompt templates drive the simulated user and the judge. Each can be replaced with your own Jinja2 file.

Override a prompt

Add the path under the matching section of eval/conftest.yml. Paths are relative to the project root and cannot leave it. Leave a key out to use the built-in template.
eval/conftest.yml
Start from the built-in text below, copy it into your project, and change what does not fit your domain.

Simulated user

Tells the model to stay in character as the user from simulation_context, keep messages short, move toward the goal one step at a time, and end the conversation once the goal is met or the agent has made clear it cannot help. It receives simulation_context and returns the next user message and a done flag.

Criteria judge

Scores each of your goals.criteria against the transcript and a ledger of the events the agent recorded. It receives criteria_text, transcript, and event_ledger. The template adapts its event vocabulary to the engine of the running server; this is the text a Mantle run uses.

Metrics judge

Scores the built-in quality metrics from the transcript. It receives transcript.

How a run is judged

Criteria

Your goals.criteria sentences. The judge receives the transcript, plus a ledger of the events the agent recorded: which skill activated and completed, which memory keys were written, which tools ran and whether they returned an error. It uses the ledger to catch an agent that says “your order is placed” without having called the tool that places it. Each criterion is returned with PASS or FAIL and a rationale. Criteria decide whether the run passes.

Built-in quality metrics

The judge also scores every run on the same fixed set of metrics, whether or not you wrote criteria: The four scores are averaged into bot_quality. The judge also writes a short summary of the conversation. Metrics are recorded, not enforced. A run can pass with task_completion: FAIL if all of its assertions and criteria passed. If task completion must be required for a scenario, write it as a criterion.

Run the evaluation

Two MCP tools do the work. Your coding agent calls them for you; you can also ask for either one by name. validate_scenario reads one scenario file and reports what is wrong with it: YAML that does not parse, a missing name or simulation_context, or an assertion type or field that is not in the scenario schema. Run it after you edit a file by hand. It does not check assertion keys against the engine; evaluate_agent does that when it reads the running server. evaluate_agent runs the scenarios. For each run it:
  1. Loads eval/conftest.yml and the scenario, and rejects any assertion key the running engine does not support.
  2. Opens a fresh sim-<uuid> conversation, sends /session_start, and applies initial_slots.
  3. Runs the simulated conversation against your agent.
  4. Fetches the tracker and evaluates the assertions.
  5. Sends the transcript and event ledger to the judge for criteria and metrics.
  6. Writes the result files and updates the batch summary.
Tell your coding agent what to run and how often:
run_count is how many times each scenario is simulated. It defaults to 1 and goes up to 10. The simulator is not deterministic, so a scenario that passes 3 of 3 says more than one that passes 1 of 1. Results are written as each run finishes; you do not have to wait for the batch to see the first failure.

Read the results

Each batch writes to eval/results/<timestamp>/, and earlier batches are kept. Running one scenario three times gives:
Every run has two files. run_N.txt is the report you read. run_N.json is the same run as data, with eval_passed, criteria, task_completion, latency, scale_metrics, and the assertion rows when the scenario had any. Use it when you script over results.

A run report

This is what one run_N.txt contains and what each part means.
  • overall_result is the pass rule applied to this run: every assertion and every criterion passed.
  • Each criterion has the judge’s rationale. Read it before changing anything; it tells you which turn the judge objected to.
  • Quality metrics are recorded and do not affect overall_result.
  • Each assertion is listed with the id it checked. A failed assertion carries a short detail saying what was found instead.
  • The transcript numbers every message. When the agent greets on session start, the transcript begins with agent: lines and no leading user: line.
  • Latency is reported for visibility and does not affect pass or fail. Time to first token (ttft) is recorded when the server streamed its reply.

The batch summary

summary.txt lists each scenario with its pass count, how long it took, and where its run files are, then a total across the batch and the time it took. When any run failed, a final Scenarios requiring attention list names them with their pass rate. summary.json holds the same numbers as data.

Use a failure as feedback

A failed criterion means the judge decided the handling was wrong; the rationale says where. A failed assertion names the fact that did not hold. If the agent did the right thing in words you did not expect, loosen the criterion. If the agent did the wrong thing, change the skill, its tool constraints, or its instructions, train again, and run the same scenario again.

Inspect a conversation

Every report ends with an Inspector URL. Open it to step through that conversation turn by turn and see the memory values and tracker events behind each reply, the same view you get from rasa inspect. Simulated conversations use a sim- prefix on the sender id, which sets them apart from conversations you held yourself.