Skip to main content
A scenario is one YAML file under eval/scenarios/. It describes a single test case: who the simulated user is, what state the conversation starts in, and what a successful conversation looks like.

Fields

goals can be left out entirely. The run then records a transcript and the built-in quality metrics, and no assertion or criterion can fail it. A failed simulation still fails the run, for example when the server is down or the simulator errors. Here is a basic example of a scenario file:
eval/scenarios/card_replace_damaged.yml

Generate, then edit

Ask your coding agent to write scenarios. It reads the skill definition and applies the conventions on this page. Give as much direction as you have:
Or, with more direction:
Generating and running are separate steps. After the files are written, open them and adjust the persona, update the simulation context, tighten a criterion, or add an assertion. If they already say what you want, run them as they are.

A full scenario

For example, this scenario checks that a customer can replace a damaged card:
eval/scenarios/card_replace_damaged.yml
The sections below explain each field in that file.

name

The label for this scenario in result files and in the batch summary.

simulation_context

A plain-language description of the simulated user. The simulator reads it before every turn, so it shapes the whole conversation. Any mix of these helps:
  • Behavior — how the user talks and reacts: calm, impatient, vague, gives up after one unclear answer.
  • Intent — what the user wants and how the conversation should unfold, including a stopping point.
  • Facts — what the user knows and brings: a card number, an order id, the name on the account.
  • Steps — a numbered sequence of what the user should do. The simulator uses the numbers to follow each step and keep track of where it is.
One sentence works. A paragraph gives a more targeted conversation. Describe the user; let the simulator choose the words.

setup.initial_slots

Optional. Each key is a memory value that exists before the first user turn. Memory can also be written by skills during the conversation; this field only sets the starting state. Use it for context the channel already has when the conversation starts, such as the caller’s phone number or preferred language, when that context is not what the scenario is testing. Each run starts a new conversation, so this does not resume an earlier one. Do not use it to skip a step the scenario depends on. If logging in changes what the agent does afterwards, leave it out and let the simulated user log in. Only project memory fields marked seed: true are copied into memory. They are sent with /session_start, the way a channel would send them. Name them either bare (caller_phone) or with the project prefix (project.caller_phone). Their values are type-checked against memory.yml. A field without seed: true is not copied into memory from initial_slots.

goals.criteria

Sentences the judge evaluates against the transcript after the conversation ends. Each one gets its own pass or fail with a written rationale. Use criteria for anything you would have to read the conversation to decide: a clean refusal, not asking for information the user already gave, not claiming something happened before it did. A failed criterion fails the run. How the judge works is described in How a run is judged.

goals.assertions

Yes-or-no checks against the events the agent recorded during the conversation. Use them for facts that must hold regardless of wording. A failed assertion fails the run. For example, to require that the order was submitted after the reason was recorded, and that the agent never sent the utter_cannot_help response:
eval/scenarios/card_replace_damaged.yml
The fields each key takes, how values are compared, and the rules for sequencing are in Assertions.

Next step

Configure the models and run the scenario: Configuration and results.