> ## Documentation Index
> Fetch the complete documentation index at: https://mantle.rasa.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Configuration and results

> The models and prompts behind a run, how a run is judged, how to run scenarios, and how to read the result files.

Everything you can set for a run lives in two places: `eval/conftest.yml` for
the models and prompts, and the `evaluate_agent` call for how many times each
scenario runs.

| Setting | Where | Required | Sets |
| - | - | - | - |
| `simulation.llm` | `eval/conftest.yml` | Yes | The model that plays the user |
| `evaluation.llm` | `eval/conftest.yml` | Yes | The model that judges criteria and quality metrics |
| `simulation.simulated_user_prompt` | `eval/conftest.yml` | No | Replaces the built-in simulated user prompt |
| `evaluation.criteria_judge_prompt` | `eval/conftest.yml` | No | Replaces the built-in criteria judge prompt |
| `evaluation.metrics_judge_prompt` | `eval/conftest.yml` | No | Replaces the built-in metrics judge prompt |
| `run_count` | `evaluate_agent` call | No, default `1`, up to `10` | How many times each scenario is simulated |

## Configure the models

`eval/conftest.yml` at the project root names two models. `simulation.llm`
plays the user. `evaluation.llm` is the judge that scores criteria and quality
metrics.

```yaml eval/conftest.yml theme={null}
simulation:
  llm:
    provider: openai
    model: gpt-5-mini

evaluation:
  llm:
    provider: openai
    model: gpt-5.1
```

Configure them independently. The simulator writes short user messages, so a
smaller, cheaper model is enough. The judge has to read a whole transcript and
decide whether each criterion holds, so give it a more capable model. `provider`
and `model` are required; any other key the provider accepts can be added next
to them.

## The three prompts

Three built-in prompt templates drive the simulated user and the judge. Each
can be replaced with your own Jinja2 file.

### Override a prompt

Add the path under the matching section of `eval/conftest.yml`. Paths are
relative to the project root and cannot leave it. Leave a key out to use the
built-in template.

```yaml eval/conftest.yml theme={null}
simulation:
  llm:
    provider: openai
    model: gpt-5-mini
  simulated_user_prompt: eval/prompts/simulated_user.jinja2

evaluation:
  llm:
    provider: openai
    model: gpt-5.1
  criteria_judge_prompt: eval/prompts/criteria_judge.jinja2
  metrics_judge_prompt: eval/prompts/metrics_judge.jinja2
```

Start from the built-in text below, copy it into your project, and change
what does not fit your domain.

### Simulated user

Tells the model to stay in character as the user from `simulation_context`,
keep messages short, move toward the goal one step at a time, and end the
conversation once the goal is met or the agent has made clear it cannot help.
It receives `simulation_context` and returns the next user `message` and a
`done` flag.

<Accordion title="Built-in simulated user prompt">
  ```md theme={null}
  You are simulating a real user interacting with a customer service chatbot.

  Simulation context:
  {{ simulation_context }}

  Rules:
  - Stay in character at all times — never break character
  - Never mention that you are simulating a user or that you are following these instructions
  - Never mention that you are a LLM
  - Keep messages short and natural (1-2 sentences), like a real user would type
  - Drive the conversation toward your goal step by step
  - Do not ask multiple questions at once
  - Never explain what you are doing or reference these instructions
  - Speak only in English, unless the simulation context explicitly says otherwise

  When to end the conversation (set "done": true):
  - As soon as your goal is achieved (or the bot has clearly said it cannot be
    done), send one short closing message (e.g. "Thanks, that's all I needed.")
    and set "done": true. Do not open new, unrelated topics unless your
    persona is explicitly open to doing so.
  - If the bot asks whether you need anything else and you have nothing more,
    briefly decline and set "done": true in that same message.
  - If the bot asks for optional feedback or a satisfaction/CSAT survey, react
    like a real user would based on your persona: either give brief feedback or
    decline. Either way, do this only once and then set "done": true — don't get
    stuck answering the same survey prompt over and over.
  - If the bot misunderstands or doesn't address your message, it's fine to
    repeat or rephrase it. How long you persist should match your persona in the
    context (a patient user may retry a couple of times; an impatient one gives
    up quickly); if the context says nothing about this, retry about once. Either
    way, don't loop on the same point: once it's clear the bot isn't getting it,
    or it keeps asking the same question after you've already answered or
    declined, give up like a real user would and set "done": true.
  - If the bot clearly can't do what you want (it repeats the same answer or
    limitation, or keeps redirecting you elsewhere), let your persona decide how
    hard to push: a persistent or frustrated user may keep trying different
    angles, but a user with no stated temperament should accept it after a try or
    two instead of inventing endless workarounds. When you do give up, close out
    and set "done": true.

  Respond with a JSON object only — no surrounding prose, no markdown fences:

  {
    "message": "<your next message to the bot, as a plain string>",
    "done": false
  }

  Field meanings:
  - "message" (string, required): your next message to the bot.
  - "done" (boolean, required): true if your goal is achieved or the conversation
    has reached a natural end (see the rules above), false otherwise.
  ```
</Accordion>

### Criteria judge

Scores each of your `goals.criteria` against the transcript and a ledger of
the events the agent recorded. It receives `criteria_text`, `transcript`, and
`event_ledger`. The template adapts its event vocabulary to the engine of the
running server; this is the text a Mantle run uses.

<Accordion title="Built-in criteria judge prompt">
  ```md theme={null}
  You are evaluating whether a chatbot conversation satisfied a list of specific requirements.

  Your only job is to mark each numbered criterion as `passed` or `not passed`. Do **not** judge tone or overall quality — that is evaluated separately. Be strict: a criterion passes only if the evidence clearly shows it was met. If evidence is ambiguous or missing, mark it as not passed.

  Write every `reasoning` field in **English**, even when the conversation transcript is in another language. You may quote transcript text in its original language when citing evidence.

  Criteria (numbered):
  {{ criteria_text }}

  Conversation transcript:
  {{ transcript }}

  Behind-the-scenes event ledger (turn-anchored timeline of what the assistant runtime actually did):

  {{ event_ledger }}

  How to use the event ledger:
  - The ledger is a curated, compressed timeline of `skill_*` / `ordered_block_*` lifecycle events, `memory_set` / `memory_cleared` events, `action: utter_<name>` renders, and `tool:` calls.
  - Each `[Tn]` block is anchored to the user's n-th message in the transcript. `[Tn] (events repeat identically through Tm — × N turns total)` means the same internal events fired in every one of those turns.

  - **`skill_completed: <skill_id>`** is the strongest single signal that a skill finished its full sequence. If a skill has a `skill_activated` but no matching `skill_completed`, the skill was interrupted and any "I've completed your X" claim by the bot is suspect.
  - **`skill_activated:` / `skill_completed:` / `skill_cancelled:` / `skill_interrupted:` / `skill_resumed:`** mean a skill started, finished, was cancelled, paused, or continued.
  - **`ordered_block_entered:` / `ordered_block_completed:`** (and cancelled/interrupted/resumed) are authored block ids written as `skill_id/block_id`.
  - **`memory_set:` / `memory_cleared:`** are memory writes and clears.
  - **`tool:` lines** are tool calls executed by the assistant runtime. They are the strongest direct evidence that a specific backend operation ran:
    - **`tool: <tool_name>(<args>) → ok`** — the runtime called `<tool_name>` with those arguments and it returned successfully. Use it to verify criteria about *which* tool was called and with what inputs — something the transcript alone cannot show.
    - **`tool: <tool_name>(<args>) → ok (empty)`** — the tool ran and reported success, but returned **no data**. Treat this as a **likely failure**: if the bot then states a concrete result that this tool was supposed to provide, that claim is **suspect** and a criterion asserting the data was delivered should `FAIL`.
    - **`tool: <tool_name>(<args>) → ok (reports error)`** — the tool returned successfully but its payload carries an error field. Treat this the same as `→ ERROR`.
    - **`tool: <tool_name>(<args>) → ERROR: <message>`** — the tool was called but failed. Any bot claim that depends on that tool succeeding is **suspect**: a criterion asserting the operation succeeded should `FAIL`.
    - The `tool:` lines are the **complete** record of the backend tools the runtime called. If a criterion requires a specific tool and it does **not** appear, the tool was not run — `FAIL` the criterion.
  - **`action: utter_<name>`** — a template / NLG render. It only produces visible bot text. It is **not** evidence that any backend operation happened. A bot saying "your transfer is complete" via `utter_transfer_complete` is just text — it does **not** prove money moved, an API was called, a record was written, an email was sent, or anything else.

  - **Precedence and hallucination-checking**: when a criterion asks whether some real-world action actually happened (a submission was made, a payment processed, a transfer executed, an account updated, a record filed):
    - Look for a successful `tool:` line (`→ ok`) in the ledger, and ideally a matching `skill_completed`. A `tool:` line that is `→ ERROR`, `→ ok (empty)`, or `→ ok (reports error)` does not count as the operation having succeeded.
    - If the bot text in the transcript says "I've done X" / "X is complete" / "X was successful" but the ledger shows neither a successful `tool:` line for the relevant operation, nor a `skill_completed` for the relevant skill, **the bot has hallucinated the outcome**. Any criterion claiming the action succeeded should `FAIL`, and any criterion phrased as "the assistant did not falsely claim X" should also `FAIL`.
    - Do **not** treat an `utter_<name>` whose name suggests success ("`utter_transfer_complete`", "`utter_invoice_submitted`") as evidence the underlying operation happened. The template can be rendered without the tool that would perform that operation having run.

  - **Events upgrade behavior verdicts; they never override a user-visible failure on an outcome criterion.** Distinguish two kinds of criteria:
    - **System-behavior criteria** — "the assistant retrieved the coupon information", "the right skill ran", "the correct tool was called", "the submission was executed". These are about what the system *did*. Use the ledger as ground truth: a `skill_completed` / successful `tool:` line can `PASS` these even if the bot's wording to the user was clumsy or generic.
    - **User-outcome / experience criteria** — "the user was able to view their coupons", "the assistant successfully helped the user reset their password", "the user received their answer". These are about what the user actually *got*. For these, the transcript is decisive: **if the bot communicated a failure or error to the user (e.g. "I'm having trouble, please try again later", "something went wrong", a generic fallback), the criterion `FAIL`s — regardless of what the ledger shows happened behind the scenes.** A backend that completed correctly while the user was shown an error is still a failed outcome for that user. Do **not** let `skill_completed` or a successful `tool:` line flip such a criterion to `PASS`.
    - When a criterion is compound ("retrieved X **and** presented it to the user"), it only `PASS`es if **both** the behavior (ledger) **and** the user-facing delivery (transcript) succeeded. A backend success behind a user-visible error `FAIL`s the compound criterion.
  - Linguistic criteria (e.g. "the assistant clearly communicated why it could not process the case") still come from the transcript — events don't have wording.
  - Do not invent events. If a criterion asks about an action or skill that does not appear in the ledger, treat that as evidence the action did not happen.

  For each criterion, echo the criterion text **verbatim** in `criterion_text` — this is used to verify index-to-verdict alignment during review.

  Respond with a JSON object only — no surrounding prose, no markdown fences:

  {
    "criteria": [
      {
        "criterion_index": 1,
        "criterion_text": "<verbatim copy of the criterion text>",
        "passed": true,
        "reasoning": "<one sentence quoting or referencing the relevant turn(s) or event(s)>"
      }
    ]
  }
  ```
</Accordion>

### Metrics judge

Scores the built-in quality metrics from the transcript. It receives
`transcript`.

<Accordion title="Built-in metrics judge prompt">
  ```md theme={null}
  You are an expert evaluator of conversational AI quality. Score the assistant's overall conversational quality along two groups of metrics:

  * **scale_metrics** — graded quality dimensions, each scored on a 1–5 scale (5 = excellent). Most production-quality conversations should land at 3–4; reserve `5` for clearly excellent behavior.
  * **binary_metrics** — pass/fail verdicts. `true` for pass, `false` for fail. No partial credit.

  Scoring rules (apply to every metric):

  1. Score the **assistant's** behavior. The user's turns are context — they tell you what the assistant should have done — but the assistant's responses are what receive the score.
  2. **Behavioral dimensions** (`helpfulness`, `repair_quality`, `coherence`) are evaluated *relative to* the user's turns: use the user's intent, corrections, and pushback to decide what the assistant should have done, then score how well the assistant actually did it.
  3. **Surface dimensions** (`tone`) score the assistant's wording itself, independent of the user's emotional state. Whether the assistant should have softened or adapted its language given user frustration is captured by `repair_quality` and `coherence`, not `tone`.
  4. Do not let the user's mood, politeness, length, or word choice move scores for the assistant. Identical assistant responses should receive identical scores regardless of how the user phrases their turns.
  5. Write all `reasoning` fields and the `summary` in **English**, even when the conversation transcript is in another language. You may quote transcript text in its original language when citing evidence.

  Conversation transcript:
  {{ transcript }}

  Binary metrics:

  - **task_completion** — was the user's intended outcome clearly delivered by the assistant?
    - `true` = the outcome was clearly delivered in the assistant's responses, not merely promised, deferred, or alluded to.
    - `false` = the outcome was missing, only partially delivered, wrong, contradicted by a later turn, the conversation ended before it was addressed, or it is unclear from the transcript whether the user got what they came for (ambiguity = `false`).

    Three rules override a naive reading:
    - **CSAT / feedback loops** after a clear outcome delivery do not change the verdict. If the outcome was delivered earlier in the conversation, return `true` regardless of subsequent feedback loops or survey prompts.
    - **Promised handoffs and async actions** — distinguish two cases:
      - **Handoff to an external party** (live agent transfer, specialist callback, escalation to a human team, "I'll have someone reach out") requires follow-through. The bot offering a handoff that the user accepts only counts as outcome delivery if the assistant actually executed it within the conversation (filed a ticket, scheduled the callback, performed the transfer, displayed a confirmation that the handoff is in motion). An offer the user accepted but the assistant never followed through on = `false`.
      - **Async system actions the assistant owns end-to-end** (email an invoice/receipt, queue a notification, submit a request to a backend, place an order, file a damage report) count as success once the assistant has collected the required inputs, confirmed them with the user, and announced the terminal action with concrete details — e.g. "your invoice will be sent to the email on file within 24 hours", "your damage report has been received". The judge cannot verify the email arrives or the report is processed in a backend — that is expected and acceptable. The verbalized completion *is* the bot's terminal action. Score `false` only if the announcement is vague or qualified ("we will try to send", "someone may follow up", "this will be handled at some point"), if the bot interrupted itself before completing the in-flow steps, or if the user clearly never confirmed the inputs.
    - **User-initiated withdrawal** counts as success only when the user gives a clear, *positive* reason to withdraw that is independent of bot performance — examples: "I already paid it", "I found it in my email", "I figured it out myself", "I changed my mind, I'll just use the app", "I don't need that anymore". The user must explicitly signal that the original need has evaporated (already met elsewhere, no longer relevant, or actively reconsidered). Only in this case return `true`.

      Two anti-patterns that are still `false`:
      - **Frustration-driven dropoffs**: short curt exits like "never mind", "forget it", "I give up", "whatever" after the bot has deflected, re-asked, ignored pushback, or otherwise failed to make progress. These signal the user gave up *because the bot failed them*, not because their need was satisfied. Without an explicit positive reason ("I found it elsewhere", "I changed my mind"), default to `false`.
      - **Pause-to-return-later**: the user steps away due to an external interruption ("hold on, I have to take a call", "I'll try again tonight", "my battery is dying"). The original need is unresolved, only deferred. Verdict is `false`.

      When in doubt about whether a withdrawal is voluntary-satisfied vs. frustration-driven, default to `false` (consistent with the general ambiguity rule).

  Scale metrics (1–5). Every point on the scale is anchored — use the in-between points (2 and 4) when the assistant's behavior clearly sits between the adjacent anchors, not as a hedge.

  The four scale metrics measure **distinct observable behaviors**. A single failure mode (e.g. a "bot stuck in a feedback loop") will normally be the failure case for *at most one* metric — pick the metric that most directly describes what went wrong, and be reluctant to penalize the others for the same observation. The "Ignore" line under each metric tells you which failure modes belong to other metrics and should not affect the current score.

  - **helpfulness** — **content quality of the assistant's substantive answers**: when the assistant attempts to answer the user's request, are those answers accurate, complete, and directly useful?
    - **Ignore**: how the assistant handled disfluency or pushback (that is `repair_quality`); whether the assistant remembered prior turns or re-asked for given info (that is `coherence`); how the answer was phrased (that is `tone`). Only the *content* of the answers themselves goes into this score.
    - **Slot-filling counts as forward progress**, not as a missing answer: an assistant that legitimately needs a piece of information to fulfill the request (e.g. asking for a phone number to look up a bill) is making productive progress and should not be scored down for "no substantive answer". Only penalize when the assistant *should* have answered something concrete and instead deflected, gave a vague non-answer, or was wrong.
    - 5 = every substantive answer was accurate, complete, and directly actionable for the user's request.
    - 4 = nearly all substantive answers were correct; one minor incomplete or imprecise answer that the user could still act on.
    - 3 = at least one significant answer was vague, generic, or only partially correct where a specific answer was needed.
    - 2 = multiple answers were unhelpfully vague or only partially correct; the user got partial value at best.
    - 1 = the assistant consistently failed to give substantive answers — only deflections, generic non-answers, or wrong information. (Score 1 even for polite, consistent deflection if the user's request was within the assistant's stated scope.)

  - **repair_quality** — **the assistant's response to user disfluency**: when the user explicitly signals the assistant misunderstood (using words like "no", "wait", "actually", "that's not what I meant", "I already told you", or repeating themselves), did the assistant adjust?
    - **Ignore**: whether the assistant's substantive answers were correct (that is `helpfulness`); whether the assistant remembered slot values (that is `coherence`); how responses were phrased (that is `tone`). This metric is *only* about behavior immediately after a user disfluency signal.
    - **If the conversation contains no disfluency signals** (the user accepted every answer and never pushed back), score 5 — there was no repair opportunity to mishandle. Note this default so it doesn't get confused with active recovery.
    - 5 = every disfluency was acknowledged and acted on; the assistant absorbed user-driven course corrections cleanly. (Or: no disfluency arose.)
    - 4 = nearly every disfluency handled; one case where the assistant needed an extra turn before recovering on its own.
    - 3 = at least one disfluency signal was missed and the user had to repeat themselves once.
    - 2 = multiple disfluencies missed; the user pushed back several times before the assistant changed course (or it never did).
    - 1 = the assistant ignored explicit user pushback or kept repeating the same response after the user said it didn't help.

  - **coherence** — **state and context tracking across turns**: does the assistant remember what was said earlier in the same conversation — slot values the user provided, prior corrections, the user's current intent?
    - **Ignore**: whether the assistant's answers were correct (that is `helpfulness`); whether it recovered from disfluency (that is `repair_quality`); how it phrased things (that is `tone`). A consistently wrong, polite deflection is highly coherent — score it accordingly. Coherence is *only* about within-conversation memory and consistency.
    - Penalize **only** for: (a) re-asking for information the user already provided, (b) contradicting an earlier statement the assistant itself made, or (c) continuing on an outdated interpretation after the user corrected it.
    - 5 = no state-tracking issues: the assistant never re-asks for given info, never contradicts itself, never reverts to an outdated user intent.
    - 4 = at most one minor slip (e.g. a brief outdated reference) that the assistant corrected within the next turn.
    - 3 = one clear state-tracking failure: re-asked for one piece of info, OR briefly followed an outdated interpretation, OR one self-contradiction.
    - 2 = two or three state-tracking failures across the conversation.
    - 1 = persistent state-tracking failure: repeatedly re-asks for already-given info, persistent contradictions, or follows an outdated interpretation across multiple turns even after the user corrected it.

  - **tone** — was the assistant's register and phrasing natural, varied, and easy to read? Do not penalize length by itself; penalize stilted, robotic, archaic, or awkward wording.
    - 5 = natural, conversational phrasing; varied and easy to read.
    - 4 = mostly natural with one or two slightly formulaic or stiff phrasings that don't disrupt readability.
    - 3 = generally fine but occasional stilted, repetitive, or awkward phrasing.
    - 2 = noticeably stilted or formulaic — multiple awkward, robotic, or repetitive phrasings that visibly affect readability.
    - 1 = robotic, archaic, or consistently stilted/awkward phrasing.

  Respond with a JSON object only — no surrounding prose, no markdown fences:

  {
    "binary_metrics": {
      "task_completion": {"passed": <true or false>, "reasoning": "<one sentence>"}
    },
    "scale_metrics": {
      "helpfulness":    {"score": <1-5>, "reasoning": "<one sentence>"},
      "repair_quality": {"score": <1-5>, "reasoning": "<one sentence>"},
      "coherence":      {"score": <1-5>, "reasoning": "<one sentence>"},
      "tone":           {"score": <1-5>, "reasoning": "<one sentence>"}
    },
    "summary": "<2-3 sentence overall summary covering strengths and main issues>"
  }
  ```
</Accordion>

## How a run is judged

### Criteria

Your `goals.criteria` sentences. The judge receives the transcript, plus a
ledger of the events the agent recorded: which skill activated and completed,
which memory keys were written, which tools ran and whether they returned an
error. It uses the ledger to catch an agent that says "your order is placed"
without having called the tool that places it. Each criterion is returned with
`PASS` or `FAIL` and a rationale. Criteria decide whether the run passes.

### Built-in quality metrics

The judge also scores every run on the same fixed set of metrics, whether or
not you wrote criteria:

| Metric | Scale | Measures |
| - | - | - |
| `task_completion` | pass or fail | Whether the user's goal was delivered in the conversation, not only promised |
| `helpfulness` | 1–5 | Whether the agent's answers were accurate, complete, and usable |
| `repair_quality` | 1–5 | Whether the agent adjusted when the user pushed back or corrected it |
| `coherence` | 1–5 | Whether the agent kept track of what was already said |
| `tone` | 1–5 | Whether the wording was natural and readable, not stilted |

The four scores are averaged into `bot_quality`. The judge also writes a short
summary of the conversation.

Metrics are recorded, not enforced. A run can pass with `task_completion:
FAIL` if all of its assertions and criteria passed. If task completion must be
required for a scenario, write it as a criterion.

## Run the evaluation

Two MCP tools do the work. Your coding agent calls them for you; you can also
ask for either one by name.

**`validate_scenario`** reads one scenario file and reports what is wrong with
it: YAML that does not parse, a missing `name` or `simulation_context`, or an
assertion type or field that is not in the scenario schema. Run it after you
edit a file by hand. It does not check assertion keys against the engine;
`evaluate_agent` does that when it reads the running server.

**`evaluate_agent`** runs the scenarios. For each run it:

1. Loads `eval/conftest.yml` and the scenario, and rejects any assertion key
   the running engine does not support.
2. Opens a fresh `sim-<uuid>` conversation, sends `/session_start`, and
   applies `initial_slots`.
3. Runs the simulated conversation against your agent.
4. Fetches the tracker and evaluates the assertions.
5. Sends the transcript and event ledger to the judge for criteria and metrics.
6. Writes the result files and updates the batch summary.

Tell your coding agent what to run and how often:

```text theme={null}
Run eval/scenarios/card_replace_damaged.yml three times.
```

```text theme={null}
Run every scenario in eval/scenarios/.
```

`run_count` is how many times each scenario is simulated. It defaults to `1`
and goes up to `10`. The simulator is not deterministic, so a scenario that
passes 3 of 3 says more than one that passes 1 of 1. Results are written as
each run finishes; you do not have to wait for the batch to see the first
failure.

## Read the results

Each batch writes to `eval/results/<timestamp>/`, and earlier batches are kept.
Running one scenario three times gives:

```text theme={null}
eval/
  conftest.yml
  scenarios/
    card_replace_damaged.yml
  results/
    2026-10-04_10-22-00/
      summary.txt
      summary.json
      card_replace_damaged/
        run_1.txt
        run_1.json
        run_2.txt
        run_2.json
        run_3.txt
        run_3.json
```

Every run has two files. `run_N.txt` is the report you read. `run_N.json` is
the same run as data, with `eval_passed`, `criteria`, `task_completion`,
`latency`, `scale_metrics`, and the assertion rows when the scenario had any.
Use it when you script over results.

### A run report

This is what one `run_N.txt` contains and what each part means.

```text theme={null}
scenario: Customer replaces a damaged card
run: 1
timestamp: 2026-10-04T10:22:11Z
conversation_id: sim-1f6de497-8e1a-45b6-8ad4-e8d68a229c50
overall_result: FAIL

--- Quality Criteria Results ---
[PASS] The agent does not claim the order is placed until it has submitted it
  rationale: The agent read back the summary, called order_replacement, and confirmed only after the tool returned.

[FAIL] The agent does not ask for the replacement reason twice
  rationale: The agent asked what happened on turn 2 and again on turn 4 after the user had answered.

--- Quality Metrics ---
bot_quality: 4.0/5
  helpfulness: 4/5 — The agent explained the replacement and the shipping options.
  repair_quality: 5/5 — The user never had to push back.
  coherence: 3/5 — The agent re-asked for the reason the user had already given.
  tone: 4/5 — Polite and direct throughout.
  task_completion: PASS — The replacement order was submitted and confirmed with a reference number.
summary: The agent completed the order but repeated one question.

--- Assertion Results ---
[PASS] skill_completed(card_replace)
[PASS] memory_was_set(card_replace.replacement_reason='damaged')
[PASS] tool_executed(order_replacement)

--- Raw Transcript ---
[1] agent: Hi, how can I help you today?
[2] user: My card has a cracked chip, I need a new one.
[3] agent: I can order a replacement. What happened to the card?
[4] user: It's damaged, the chip is cracked.
...

--- Latency ---
turn 1: ttft=612ms  turn_latency=1420ms
turn 2: ttft=580ms  turn_latency=1310ms
conversation avg: ttft=596ms  turn_latency=1365ms  (over 2 turns with bot text)

--- Inspector URL ---
http://localhost:5005/webhooks/inspector/inspect.html?sender=sim-1f6de497-8e1a-45b6-8ad4-e8d68a229c50
```

* `overall_result` is the pass rule applied to this run: every assertion and
  every criterion passed.
* Each criterion has the judge's rationale. Read it before changing anything;
  it tells you which turn the judge objected to.
* Quality metrics are recorded and do not affect `overall_result`.
* Each assertion is listed with the id it checked. A failed assertion carries
  a short detail saying what was found instead.
* The transcript numbers every message. When the agent greets on session
  start, the transcript begins with `agent:` lines and no leading `user:` line.
* Latency is reported for visibility and does not affect pass or fail.
  Time to first token (`ttft`) is recorded when the server streamed its reply.

### The batch summary

`summary.txt` lists each scenario with its pass count, how long it took, and
where its run files are, then a total across the batch and the time it took.
When any run failed, a final **Scenarios requiring attention** list names them
with their pass rate. `summary.json` holds the same numbers as data.

### Use a failure as feedback

A failed criterion means the judge decided the handling was wrong; the
rationale says where. A failed assertion names the fact that did not hold. If
the agent did the right thing in words you did not expect, loosen the
criterion. If the agent did the wrong thing, change the skill, its tool
constraints, or its instructions, train again, and run the same scenario again.

## Inspect a conversation

Every report ends with an Inspector URL. Open it to step through that
conversation turn by turn and see the memory values and tracker events behind
each reply, the same view you get from `rasa inspect`.

Simulated conversations use a `sim-` prefix on the sender id, which sets them
apart from conversations you held yourself.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.