goals.assertions in a scenario file. Each one is a
yes-or-no check against the events the agent recorded during the simulated
conversation, so the result does not depend on how the agent phrased anything.
The checks are grouped by what they look at: skills, ordered blocks, memory,
tools, agent messages, and sequencing.
Unless a section says otherwise, a check passes when at least one recorded
event matches it. The event payloads these checks read are listed in
Tracker events. Where assertions
sit in a scenario file is in Scenarios.
Skill
Skill checks match the lifecycle events of one skill.skill_id is the skill
folder name, for example card_replace for skills/card_replace/skill.md. It
is not the name: from the skill’s frontmatter.
eval/scenarios/card_replace_damaged.yml
step_id narrows a cancel, interrupt, or resume check to the step the skill
was on when the event was recorded. Omit it unless you know the recorded
value: when the skill was running prose, the recorded value is autonomous;
when it was inside an ordered block, it is the id of the block step it was on.
skill_activated and skill_completed reject step_id.
An ordered-block event does not satisfy a skill check, and the other way
around.
Ordered block
Ordered-block checks match the lifecycle events of one ordered block.block_id is the id you wrote
in the skill, such as main or submit. skill_id is the skill that owns
the block.
eval/scenarios/card_replace_damaged.yml
step_id here is the step the block was on when the event was recorded.
ordered_block_entered rejects it.
Memory
Memory checks look at writes to and clears of one memory key. Thename is
the key exactly as the tracker stores it: the scope and the field, joined with
a dot. card_replace.replacement_reason is a skill field,
project.customer_authenticated a project field, and system.channel a
system field. Do not add a session. prefix; the check would never match.
Each memory key takes a list of entries. Every entry in the list must pass.
How values are compared:
- The type must match.
truedoes not match"true", and5does not match"5". - Objects match regardless of key order. Lists match only in the same order.
nullis not accepted as a value.
memory_was_set passes, memory_was_cleared passes, and
memory_was_not_set fails.
eval/scenarios/card_replace_damaged.yml
Tools
tool_executed
Passes when one recorded tool run matches every field you set.
eval/scenarios/card_replace_damaged.yml
tools_within_one_turn
Passes when every listed call happened during the same user turn. Use it when
the agent should gather several results before answering, instead of asking
the user to wait between tools.
Mantle’s own lifecycle tools are not counted in this check, except
search_knowledge and cannot_help. So complete_skill, set_fields,
listen, and hangup neither satisfy a listed call nor count as an extra
call. A standalone tool_executed check still sees those runs.
eval/scenarios/card_replace_damaged.yml
Agent messages
bot_uttered passes when at least one agent message matches.
bot_did_not_utter passes when no agent message matches. Set at least one of:
eval/scenarios/card_replace_damaged.yml
Sequencing
sequencing takes a list of checks and passes when each one matches an event
that comes after the event the previous check matched. Other events may sit in
between; the next check does not have to match the very next event. Use it
when the order matters, for example that a tool ran only after a memory value
was recorded.
eval/scenarios/card_replace_damaged.yml
sequencing:
memory_was_setandmemory_was_clearedtake the key name as a string, not a list, and cannot carry avalue.- Skill, ordered-block, and
tool_executedchecks use the same object shape as above. memory_was_not_setandtools_within_one_turnare not allowed.
See also
- Scenarios: the file these assertions live in
- Configuration and results: run a scenario and read the result
- Tracker events: the events these checks read
- Memory: key names and
seed: trueproject fields