> ## Documentation Index
> Fetch the complete documentation index at: https://mantle.rasa.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Assertions

> Deterministic checks a Mantle evaluation scenario can make against tracker events.

Assertions live under `goals.assertions` in a scenario file. Each one is a
yes-or-no check against the events the agent recorded during the simulated
conversation, so the result does not depend on how the agent phrased anything.
The checks are grouped by what they look at: skills, ordered blocks, memory,
tools, agent messages, and sequencing.

Unless a section says otherwise, a check passes when at least one recorded
event matches it. The event payloads these checks read are listed in
[Tracker events](/reference/execution-loop#tracker-events). Where assertions
sit in a scenario file is in [Scenarios](/testing/scenarios).

## Skill

Skill checks match the lifecycle events of one skill. `skill_id` is the skill
folder name, for example `card_replace` for `skills/card_replace/skill.md`. It
is not the `name:` from the skill's frontmatter.

| Key | Required | Optional |
| - | - | - |
| `skill_activated` | `skill_id` | — |
| `skill_completed` | `skill_id` | — |
| `skill_cancelled` | `skill_id` | `step_id` |
| `skill_interrupted` | `skill_id` | `step_id` |
| `skill_resumed` | `skill_id` | `step_id` |

```yaml eval/scenarios/card_replace_damaged.yml theme={null}
assertions:
  - skill_activated:
      skill_id: card_replace
  - skill_completed:
      skill_id: card_replace
```

`step_id` narrows a cancel, interrupt, or resume check to the step the skill
was on when the event was recorded. Omit it unless you know the recorded
value: when the skill was running prose, the recorded value is `autonomous`;
when it was inside an ordered block, it is the id of the block step it was on.
`skill_activated` and `skill_completed` reject `step_id`.

An ordered-block event does not satisfy a skill check, and the other way
around.

## Ordered block

Ordered-block checks match the lifecycle events of one
[ordered block](/build-guide/ordered-blocks). `block_id` is the id you wrote
in the skill, such as `main` or `submit`. `skill_id` is the skill that owns
the block.

| Key | Required | Optional |
| - | - | - |
| `ordered_block_entered` | `skill_id`, `block_id` | — |
| `ordered_block_completed` | `skill_id`, `block_id` | `step_id` |
| `ordered_block_cancelled` | `skill_id`, `block_id` | `step_id` |
| `ordered_block_interrupted` | `skill_id`, `block_id` | `step_id` |
| `ordered_block_resumed` | `skill_id`, `block_id` | `step_id` |

```yaml eval/scenarios/card_replace_damaged.yml theme={null}
assertions:
  - ordered_block_entered:
      skill_id: card_replace
      block_id: main
```

`step_id` here is the step the block was on when the event was recorded.
`ordered_block_entered` rejects it.

## Memory

Memory checks look at writes to and clears of one memory key. The `name` is
the key exactly as the tracker stores it: the scope and the field, joined with
a dot. `card_replace.replacement_reason` is a skill field,
`project.customer_authenticated` a project field, and `system.channel` a
system field. Do not add a `session.` prefix; the check would never match.

Each memory key takes a list of entries. Every entry in the list must pass.

| Key | `name` | `value` |
| - | - | - |
| `memory_was_set` | Required | Optional. Without it, any write to the key passes. With it, the written value must equal it exactly. |
| `memory_was_not_set` | Required | Optional. Without it, the check fails if the key was written at all. With it, the check fails only if the last write to the key was that value. |
| `memory_was_cleared` | Required | Do not set one. A clear has no value. The check passes when that key was cleared. |

How values are compared:

* The type must match. `true` does not match `"true"`, and `5` does not match
  `"5"`.
* Objects match regardless of key order. Lists match only in the same order.
* `null` is not accepted as a value.

Writes and clears are separate events. If the key is written and later
cleared, `memory_was_set` passes, `memory_was_cleared` passes, and
`memory_was_not_set` fails.

```yaml eval/scenarios/card_replace_damaged.yml theme={null}
assertions:
  - memory_was_set:
      - name: card_replace.replacement_reason
        value: damaged
      - name: card_replace.selected_card_id
  - memory_was_not_set:
      - name: card_replace.fraud_review
  - memory_was_cleared:
      - name: card_replace.draft_order_id
```

## Tools

### `tool_executed`

Passes when one recorded tool run matches every field you set.

| Field | Required | Notes |
| - | - | - |
| `tool_name` | Yes | |
| `source` | Yes | `local` or `mcp`. A local run does not match an MCP run of the same name. |
| `mcp_server` | When `source` is `mcp` | Not allowed when `source` is `local`. It is not inferred from the tool name. |
| `skill_id` | No | When set, the run must have been made on behalf of that skill. Omit it to match a run from any skill, including none. |
| `arguments` | No | Each listed key must equal the value the tool was called with. Arguments you do not list are ignored. |
| `is_error` | No | `true` matches only failed runs, `false` only successful ones. Omit it to match both. |

```yaml eval/scenarios/card_replace_damaged.yml theme={null}
assertions:
  - tool_executed:
      tool_name: order_replacement
      source: local
      skill_id: card_replace
      arguments:
        shipping: standard
  - tool_executed:
      tool_name: submit_transfer
      source: mcp
      mcp_server: banking
```

### `tools_within_one_turn`

Passes when every listed call happened during the same user turn. Use it when
the agent should gather several results before answering, instead of asking
the user to wait between tools.

| Field | Required | Notes |
| - | - | - |
| `calls` | Yes | At least two `tool_executed` entries, in the shape above. |
| `allow_extra_calls` | Yes | `false` fails the turn if it contains any other counted tool run. `true` allows them. |
| `ordered` | No | Defaults to `false`. `true` requires the listed order. `false` accepts any order; each tool run counts toward only one listed call. |

Mantle's own lifecycle tools are not counted in this check, except
`search_knowledge` and `cannot_help`. So `complete_skill`, `set_fields`,
`listen`, and `hangup` neither satisfy a listed call nor count as an extra
call. A standalone `tool_executed` check still sees those runs.

```yaml eval/scenarios/card_replace_damaged.yml theme={null}
assertions:
  - tools_within_one_turn:
      ordered: true
      allow_extra_calls: false
      calls:
        - tool_executed:
            tool_name: check_balance
            source: local
        - tool_executed:
            tool_name: order_replacement
            source: local
```

## Agent messages

`bot_uttered` passes when at least one agent message matches.
`bot_did_not_utter` passes when no agent message matches. Set at least one of:

| Field | Matches |
| - | - |
| `utter_name` | The response name the message was rendered from. |
| `text_matches` | A regular expression searched anywhere in the message text. |
| `buttons` | A list of `title` and optional `payload` the message carried. |

```yaml eval/scenarios/card_replace_damaged.yml theme={null}
assertions:
  - bot_uttered:
      text_matches: "replacement.*(ordered|placed)"
  - bot_did_not_utter:
      utter_name: utter_cannot_help
```

## Sequencing

`sequencing` takes a list of checks and passes when each one matches an event
that comes after the event the previous check matched. Other events may sit in
between; the next check does not have to match the very next event. Use it
when the order matters, for example that a tool ran only after a memory value
was recorded.

```yaml eval/scenarios/card_replace_damaged.yml theme={null}
assertions:
  - sequencing:
      - memory_was_set: card_replace.replacement_reason
      - tool_executed:
          tool_name: order_replacement
          source: local
      - skill_completed:
          skill_id: card_replace
```

Inside `sequencing`:

* `memory_was_set` and `memory_was_cleared` take the key name as a string, not
  a list, and cannot carry a `value`.
* Skill, ordered-block, and `tool_executed` checks use the same object shape as
  above.
* `memory_was_not_set` and `tools_within_one_turn` are not allowed.

## See also

* [Scenarios](/testing/scenarios): the file these assertions live in
* [Configuration and results](/testing/configuration-and-results): run a scenario and read the result
* [Tracker events](/reference/execution-loop#tracker-events): the events these checks read
* [Memory](/reference/memory-yml): key names and `seed: true` project fields


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.