> ## Documentation Index
> Fetch the complete documentation index at: https://mantle.rasa.com/llms.txt
> Use this file to discover all available pages before exploring further.

# The Runtime Loop

> What Mantle does on every turn, step by step.

Every turn runs the same loop, and the order is the point: the **framework acts
first**, and only calls the LLM for what genuinely needs judgment. That's what
"deterministic-first" means.

## Each turn, in order

1. **The framework acts first.** It applies an inbound
   [`set:` button tap](/reference/responses-yml#set-memory-payloads) when
   present, then runs whatever is already determined, such as
   the next steps in an active [ordered block](/build-guide/ordered-blocks)
   (including a `collect:` whose value was just written),
   without calling the LLM. It also narrows what the LLM may do next: which
   skills are routable (`hidden` / `disabled` / engine-managed are omitted) and which tools are visible.
2. **It builds a focused prompt.** Only the active skill's instructions, the
   tools currently allowed, and the memory values that matter. See
   [Context Management](/mantle/context).
3. **It calls the LLM.** The model replies, calls a tool, or switches skills,
   but only within what the framework already allowed. Choosing a skill is the
   model's call; the framework decides which skills were eligible to choose
   from. It may also call `listen` to wait without sending a message.
4. **Tool calls run back through the framework.** Constraints and confirmations
   are checked before the tool runs, then the loop repeats from step 1.

One user message can loop several times. Call a tool, settle the result, call
another. The framework is in the path every time, not just at the start.

After a canned ordered-block reply, leftover hybrid prose can continue on the
same turn instead of waiting for the next user message. See
[When the engine waits](/reference/execution-loop#when-the-engine-waits).

```mermaid theme={null}
flowchart TD
    A[User message] --> B[Framework runs first]
    B --> C{Can the framework<br/>settle this itself?}
    C -->|Yes| D[Act without the LLM]
    D --> J{Wait for the user?}
    C -->|No. Needs judgment| E[Build a focused prompt]
    E --> F[Call the LLM]
    F --> G{What did it do?}
    G -->|Called a tool| H[Run the tool]
    H --> B
    G -->|Replied, asked, or listen| I[Send message if any]
    I --> J
    J -->|Empty stack, collect, listen| K[Wait]
    J -->|More work remains| B
```

## Why deterministic-first

Anything the framework can settle on its own is faster, cheaper, and can't be
talked out of. The next deterministic step in an
[ordered block](/build-guide/ordered-blocks) is resolved without a model call.
The LLM is reserved for what actually needs reasoning, and even then it acts
only inside the boundaries the framework has already set. See
[Guarantees & Guardrails](/mantle/guarantees).

## Running tools

A skill can call two kinds of tools: its own, auto-discovered from the skill's
`tools.py` or `tools/` folder, and shared tools it imports from the agent root.
On a name collision the skill-local tool wins. This precedence is resolved when
the model loads, not looked up per call.

When the LLM calls a tool, its [constraints](/build-guide/tool-constraints) are
checked first: a tool whose `requires:` condition is unmet is never offered to
the LLM at all. A tool can also write to memory as it runs, so acting and
recording happen in one step.

The same function can be run directly by an ordered-block `execute_tool:` step.
The block controls when the step is reached, and the same `requires:` gate is
checked before the function runs. See [Tools](/concepts/tools) for the full
contract. Successful and failed dispatches are recorded as `tool_executed`. See
[Tracker events](/reference/execution-loop#tracker-events).

## Start a conversation

The bundled `default_session_start` skill initializes a session without speaking.
How you customize it depends on who talks first. The other bundled skills and
their default wording are in [Default skills](/reference/default-skills).

### The agent speaks first

Use this when the channel should greet before the user types or speaks. Override
`default_session_start` with a greeting, and have the frontend send `/session_start`
as the first user message so that greeting runs on its own turn.

`rasa shell`, the Inspector, and the built-in voice channels already send
`/session_start`. A custom REST or web client must do the same.

```markdown skills/default_session_start/skill.md theme={null}
---
name: Session Start
description: Greet the user at the start of a session.
routing:
  engine_managed: true
---
:::ordered_block id=main
steps:
  - id: greet
    action: utter_greet
:::
```

```yaml skills/default_session_start/responses.yml theme={null}
responses:
  utter_greet:
    - text: "Thanks for calling Telco support. How can I help?"
```

Keep `routing.engine_managed: true` so the skill stays engine-activated rather
than model-routable.

### The user speaks first

Use this when the first request is a real user utterance (typical REST clients
that do not send `/session_start`). Leave the bundled skill as it is — no greeting
override. The engine starts the session, then handles that first message in the
same request.

A greeting override on this path still runs before the user's request is handled,
so the user hears a welcome and then an answer in one turn. Skip the greeting
until the channel can send `/session_start` first.

## When a turn is cancelled

A turn can end before every in-flight step finishes, whether from a voice
barge-in, a session timeout, or a dropped connection. The runtime races
cancellation against LLM
calls, tool dispatch, and skill advancement: when cancel wins, in-flight work is
aborted promptly (with a short grace window for teardown) and the turn returns
as cancelled rather than as an error.

For tool authors, the important part is what happens inside a running tool:
`asyncio.CancelledError` at the current `await`, no call record if the tool was
mid-`await`, memory already written is kept, and the step can run again next
turn. See [Cancellation](/reference/tools#cancellation) for `context.is_cancelled`,
cleanup, and idempotency patterns.

## Reference

For the loop expressed as a terse spec, plus what state the framework tracks at
each control level, see the [Execution Loop reference](/reference/execution-loop).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.