- Assertions are yes-or-no checks against the events the agent recorded: a skill completed, a memory key was written, a tool ran.
- Criteria are sentences you write in plain language. An LLM judge reads the transcript and the recorded events and decides whether each one holds.
When to use
rasa train checks that your project is valid and packages it into a model.
rasa inspect lets you talk to the agent yourself. Neither tells you how the
agent handles many conversations, or whether a change you made broke a path
you did not try.
Use it while you are building the agent, and again after the agent is in
production when you extend its use cases. The runs check that the existing
use cases still work, and that the agent still does what your scenarios
require. Write the scenarios for the agent as a whole, not only for the skill
you are editing.
Before you start
You need four things in place.- A running agent. Start it with
rasa inspect, or withrasa run --inspect. Both turn on the conversations API the simulator uses to read the tracker and applyinitial_slots. - The
restandinspectorchannels enabled inintegrations.yml. The simulator talks to the agent overrest.inspectorserves the per-run link you use to step through a conversation afterwards. - The Mantle skills for your coding agent and the Rasa MCP server
(
rasa tools run), set up as in Getting started. Themantle-simulating-evaluatingskill in that pack tells your coding agent how to write scenario files for a Mantle project and which MCP tools to call. Two tools matter here:validate_scenarioandevaluate_agent. - API keys for the LLM providers you use, set as in
Getting started. The
simulated user and the judge are LLM calls, and their models are named in
eval/conftest.yml(see Configuration and results).
Quick start
Createeval/conftest.yml once, as shown in
Configure the models.
Then, with the agent running, ask your coding agent for a scenario:
- Reads the skill’s
skill.md,memory.yml, and tools to learn what the agent can do and which events it records. - Writes one scenario file per case under
eval/scenarios/. - Calls
validate_scenarioon each file. - Calls
evaluate_agent, waits for the runs, and reports pass or fail for each one with a link to the result files.
Next
- Scenarios: the file that describes one simulated conversation and its checks
- Configuration and results: the models and prompts, how a run is judged, and how to read the result files