eval/scenarios/. It describes a single
test case: who the simulated user is, what state the conversation starts in,
and what a successful conversation looks like.
Fields
goals can be left out entirely. The run then records a transcript and the
built-in quality metrics, and no assertion or criterion can fail it. A failed
simulation still fails the run, for example when the server is down or the
simulator errors.
Here is a basic example of a scenario file:
eval/scenarios/card_replace_damaged.yml
Generate, then edit
Ask your coding agent to write scenarios. It reads the skill definition and applies the conventions on this page. Give as much direction as you have:A full scenario
For example, this scenario checks that a customer can replace a damaged card:eval/scenarios/card_replace_damaged.yml
name
The label for this scenario in result files and in the batch summary.
simulation_context
A plain-language description of the simulated user. The simulator reads it
before every turn, so it shapes the whole conversation. Any mix of these helps:
- Behavior — how the user talks and reacts: calm, impatient, vague, gives up after one unclear answer.
- Intent — what the user wants and how the conversation should unfold, including a stopping point.
- Facts — what the user knows and brings: a card number, an order id, the name on the account.
- Steps — a numbered sequence of what the user should do. The simulator uses the numbers to follow each step and keep track of where it is.
setup.initial_slots
Optional. Each key is a memory value that exists before the first user turn.
Memory can also be written by skills during the conversation; this field only
sets the starting state.
Use it for context the channel already has when the conversation starts, such
as the caller’s phone number or preferred language, when that context is not
what the scenario is testing. Each run starts a new conversation, so this does
not resume an earlier one. Do not use it to skip a step the scenario depends
on. If logging in changes what the agent does afterwards, leave it out and let
the simulated user log in.
Only project memory fields marked
seed: true are copied
into memory. They are sent with /session_start, the way a channel would send
them. Name them either bare (caller_phone) or with the project prefix
(project.caller_phone). Their values are type-checked against memory.yml.
A field without seed: true is not copied into memory from initial_slots.
goals.criteria
Sentences the judge evaluates against the transcript after the conversation
ends. Each one gets its own pass or fail with a written rationale. Use criteria
for anything you would have to read the conversation to decide: a clean
refusal, not asking for information the user already gave, not claiming
something happened before it did.
A failed criterion fails the run. How the judge works is described in
How a run is judged.
goals.assertions
Yes-or-no checks against the events the agent recorded during the
conversation. Use them for facts that must hold regardless of wording. A
failed assertion fails the run.
For example, to require that the order was submitted after the reason was
recorded, and that the agent never sent the
utter_cannot_help response:
eval/scenarios/card_replace_damaged.yml
sequencing are in Assertions.