Skip to main content
Simulation and evaluation tests how your agent behaves in conversations you did not type yourself. You set the context for a simulated conversation by describing how the simulated user should behave and what the goal of the conversation is. An LLM plays that user and holds a full conversation with your running agent. When the conversation ends, two kinds of checks run against it:
  • Assertions are yes-or-no checks against the events the agent recorded: a skill completed, a memory key was written, a tool ran.
  • Criteria are sentences you write in plain language. An LLM judge reads the transcript and the recorded events and decides whether each one holds.
One conversation with its checks is a run. A run passes when every assertion and every criterion passes. The judge also scores a fixed set of quality metrics on every run. The scores are saved with the result. Only the assertions and criteria decide whether the run passes.

When to use

rasa train checks that your project is valid and packages it into a model. rasa inspect lets you talk to the agent yourself. Neither tells you how the agent handles many conversations, or whether a change you made broke a path you did not try. Use it while you are building the agent, and again after the agent is in production when you extend its use cases. The runs check that the existing use cases still work, and that the agent still does what your scenarios require. Write the scenarios for the agent as a whole, not only for the skill you are editing.

Before you start

You need four things in place.
  1. A running agent. Start it with rasa inspect, or with rasa run --inspect. Both turn on the conversations API the simulator uses to read the tracker and apply initial_slots.
  2. The rest and inspector channels enabled in integrations.yml. The simulator talks to the agent over rest. inspector serves the per-run link you use to step through a conversation afterwards.
  3. The Mantle skills for your coding agent and the Rasa MCP server (rasa tools run), set up as in Getting started. The mantle-simulating-evaluating skill in that pack tells your coding agent how to write scenario files for a Mantle project and which MCP tools to call. Two tools matter here: validate_scenario and evaluate_agent.
  4. API keys for the LLM providers you use, set as in Getting started. The simulated user and the judge are LLM calls, and their models are named in eval/conftest.yml (see Configuration and results).

Quick start

Create eval/conftest.yml once, as shown in Configure the models. Then, with the agent running, ask your coding agent for a scenario:
Or for several at once:
The coding agent then:
  1. Reads the skill’s skill.md, memory.yml, and tools to learn what the agent can do and which events it records.
  2. Writes one scenario file per case under eval/scenarios/.
  3. Calls validate_scenario on each file.
  4. Calls evaluate_agent, waits for the runs, and reports pass or fail for each one with a link to the result files.
Open the files it wrote. Edit the persona, tighten a criterion, or add an assertion, or run them as they are if they already say what you want. The Scenarios page explains every field.

Next

  • Scenarios: the file that describes one simulated conversation and its checks
  • Configuration and results: the models and prompts, how a run is judged, and how to read the result files