> ## Documentation Index
> Fetch the complete documentation index at: https://mantle.rasa.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Scenarios

> The YAML file that describes one simulated conversation: who the user is, how the conversation starts, and what counts as success.

A scenario is one YAML file under `eval/scenarios/`. It describes a single
test case: who the simulated user is, what state the conversation starts in,
and what a successful conversation looks like.

## Fields

| Field | Required | Sets |
| - | - | - |
| `name` | Yes | The label for this scenario in result files and the batch summary |
| `simulation_context` | Yes | How the simulated user behaves and what they are trying to do |
| `setup.initial_slots` | No | Memory values that exist before the first user turn |
| `goals.criteria` | No | Sentences the judge scores; a failed criterion fails the run |
| `goals.assertions` | No | Checks against recorded events; a failed assertion fails the run |

`goals` can be left out entirely. The run then records a transcript and the
built-in quality metrics, and no assertion or criterion can fail it. A failed
simulation still fails the run, for example when the server is down or the
simulator errors.

Here is a basic example of a scenario file:

```yaml eval/scenarios/card_replace_damaged.yml theme={null}
scenario:
  name: Customer replaces a damaged card
  simulation_context: >
    You are a calm customer whose card has a cracked chip. Ask to replace it.
  goals:
    criteria:
      - The agent confirms the replacement only after the order is submitted
    assertions:
      - skill_completed:
          skill_id: card_replace
```

## Generate, then edit

Ask your coding agent to write scenarios. It reads the skill definition and
applies the conventions on this page. Give as much direction as you have:

```text theme={null}
Generate scenarios for the card_replace skill.
```

Or, with more direction:

```text theme={null}
Generate a scenario where an impatient customer wants a replacement for a
card with a cracked chip. Assert that card_replace completes and that the
agent never asks for the reason twice.
```

Generating and running are separate steps. After the files are written, open
them and adjust the persona, update the simulation context, tighten a
criterion, or add an assertion. If they already say what you want, run them
as they are.

## A full scenario

For example, this scenario checks that a customer can replace a damaged card:

```yaml eval/scenarios/card_replace_damaged.yml theme={null}
scenario:
  name: Customer replaces a damaged card

  simulation_context: >
    You are a calm customer whose card has a cracked chip. Ask to replace the
    card. When the agent asks what happened, say it is damaged. Confirm when
    the order summary is read back. Do not bring up anything else.

  setup:
    initial_slots:
      project.preferred_language: en

  goals:
    criteria:
      - The agent does not claim the order is placed until it has submitted it
      - The agent does not ask for the replacement reason twice
    assertions:
      - skill_completed:
          skill_id: card_replace
      - memory_was_set:
          - name: card_replace.replacement_reason
            value: damaged
      - tool_executed:
          tool_name: order_replacement
          source: local
```

The sections below explain each field in that file.

## `name`

The label for this scenario in result files and in the batch summary.

## `simulation_context`

A plain-language description of the simulated user. The simulator reads it
before every turn, so it shapes the whole conversation. Any mix of these helps:

* **Behavior** — how the user talks and reacts: calm, impatient, vague, gives
  up after one unclear answer.
* **Intent** — what the user wants and how the conversation should unfold,
  including a stopping point.
* **Facts** — what the user knows and brings: a card number, an order id, the
  name on the account.
* **Steps** — a numbered sequence of what the user should do. The simulator
  uses the numbers to follow each step and keep track of where it is.

One sentence works. A paragraph gives a more targeted conversation. Describe
the user; let the simulator choose the words.

## `setup.initial_slots`

Optional. Each key is a memory value that exists before the first user turn.
Memory can also be written by skills during the conversation; this field only
sets the starting state.

Use it for context the channel already has when the conversation starts, such
as the caller's phone number or preferred language, when that context is not
what the scenario is testing. Each run starts a new conversation, so this does
not resume an earlier one. Do not use it to skip a step the scenario depends
on. If logging in changes what the agent does afterwards, leave it out and let
the simulated user log in.

Only project memory fields marked
[`seed: true`](/reference/memory-yml#seeding-from-session-start) are copied
into memory. They are sent with `/session_start`, the way a channel would send
them. Name them either bare (`caller_phone`) or with the project prefix
(`project.caller_phone`). Their values are type-checked against `memory.yml`.
A field without `seed: true` is not copied into memory from `initial_slots`.

## `goals.criteria`

Sentences the judge evaluates against the transcript after the conversation
ends. Each one gets its own pass or fail with a written rationale. Use criteria
for anything you would have to read the conversation to decide: a clean
refusal, not asking for information the user already gave, not claiming
something happened before it did.

A failed criterion fails the run. How the judge works is described in
[How a run is judged](/testing/configuration-and-results#how-a-run-is-judged).

## `goals.assertions`

Yes-or-no checks against the events the agent recorded during the
conversation. Use them for facts that must hold regardless of wording. A
failed assertion fails the run.

| Assertion | Checks |
| - | - |
| `skill_activated` | A skill was activated |
| `skill_completed` | A skill reached completion |
| `skill_cancelled` | A skill was cancelled |
| `skill_interrupted` | A skill was paused before it finished |
| `skill_resumed` | A skill continued after an interruption |
| `ordered_block_entered` | An ordered block was entered |
| `ordered_block_completed` | An ordered block finished |
| `ordered_block_cancelled` | An ordered block was cancelled |
| `ordered_block_interrupted` | An ordered block was interrupted |
| `ordered_block_resumed` | An ordered block continued after an interruption |
| `memory_was_set` | A memory key was written, optionally to a given value |
| `memory_was_not_set` | A memory key was never written, or never to a given value |
| `memory_was_cleared` | A memory key was cleared |
| `tool_executed` | A tool ran, optionally with given arguments or on behalf of a given skill |
| `tools_within_one_turn` | Several tools ran during the same user turn |
| `bot_uttered` | The agent sent a message matching a response name, text, or buttons |
| `bot_did_not_utter` | The agent sent no message matching the pattern |
| `sequencing` | The listed checks matched events in the listed order |

For example, to require that the order was submitted after the reason was
recorded, and that the agent never sent the `utter_cannot_help` response:

```yaml eval/scenarios/card_replace_damaged.yml theme={null}
assertions:
  - sequencing:
      - memory_was_set: card_replace.replacement_reason
      - tool_executed:
          tool_name: order_replacement
          source: local
  - bot_did_not_utter:
      utter_name: utter_cannot_help
```

The fields each key takes, how values are compared, and the rules for
`sequencing` are in [Assertions](/reference/assertions).

## Next step

Configure the models and run the scenario:
[Configuration and results](/testing/configuration-and-results).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.