Skip to main content
Every turn runs the same loop, and the order is the point: the framework acts first, and only calls the LLM for what genuinely needs judgment. That’s what “deterministic-first” means.

Each turn, in order

  1. The framework acts first. It applies an inbound set: button tap when present, then runs whatever is already determined, such as the next steps in an active ordered block (including a collect: whose value was just written), without calling the LLM. It also narrows what the LLM may do next: which skills are routable (hidden / disabled / engine-managed are omitted) and which tools are visible.
  2. It builds a focused prompt. Only the active skill’s instructions, the tools currently allowed, and the memory values that matter. See Context Management.
  3. It calls the LLM. The model replies, calls a tool, or switches skills, but only within what the framework already allowed. Choosing a skill is the model’s call; the framework decides which skills were eligible to choose from. It may also call listen to wait without sending a message.
  4. Tool calls run back through the framework. Constraints and confirmations are checked before the tool runs, then the loop repeats from step 1.
One user message can loop several times. Call a tool, settle the result, call another. The framework is in the path every time, not just at the start. After a canned ordered-block reply, leftover hybrid prose can continue on the same turn instead of waiting for the next user message. See When the engine waits.

Why deterministic-first

Anything the framework can settle on its own is faster, cheaper, and can’t be talked out of. The next deterministic step in an ordered block is resolved without a model call. The LLM is reserved for what actually needs reasoning, and even then it acts only inside the boundaries the framework has already set. See Guarantees & Guardrails.

Running tools

A skill can call two kinds of tools: its own, auto-discovered from the skill’s tools.py or tools/ folder, and shared tools it imports from the agent root. On a name collision the skill-local tool wins. This precedence is resolved when the model loads, not looked up per call. When the LLM calls a tool, its constraints are checked first: a tool whose requires: condition is unmet is never offered to the LLM at all. A tool can also write to memory as it runs, so acting and recording happen in one step. The same function can be run directly by an ordered-block execute_tool: step. The block controls when the step is reached, and the same requires: gate is checked before the function runs. See Tools for the full contract. Successful and failed dispatches are recorded as tool_executed. See Tracker events.

Start a conversation

The bundled default_session_start skill initializes a session without speaking. How you customize it depends on who talks first. The other bundled skills and their default wording are in Default skills.

The agent speaks first

Use this when the channel should greet before the user types or speaks. Override default_session_start with a greeting, and have the frontend send /session_start as the first user message so that greeting runs on its own turn. rasa shell, the Inspector, and the built-in voice channels already send /session_start. A custom REST or web client must do the same.
skills/default_session_start/skill.md
skills/default_session_start/responses.yml
Keep routing.engine_managed: true so the skill stays engine-activated rather than model-routable.

The user speaks first

Use this when the first request is a real user utterance (typical REST clients that do not send /session_start). Leave the bundled skill as it is — no greeting override. The engine starts the session, then handles that first message in the same request. A greeting override on this path still runs before the user’s request is handled, so the user hears a welcome and then an answer in one turn. Skip the greeting until the channel can send /session_start first.

When a turn is cancelled

A turn can end before every in-flight step finishes, whether from a voice barge-in, a session timeout, or a dropped connection. The runtime races cancellation against LLM calls, tool dispatch, and skill advancement: when cancel wins, in-flight work is aborted promptly (with a short grace window for teardown) and the turn returns as cancelled rather than as an error. For tool authors, the important part is what happens inside a running tool: asyncio.CancelledError at the current await, no call record if the tool was mid-await, memory already written is kept, and the step can run again next turn. See Cancellation for context.is_cancelled, cleanup, and idempotency patterns.

Reference

For the loop expressed as a terse spec, plus what state the framework tracks at each control level, see the Execution Loop reference.