set_fields fill. On those three
calls, RetryModel discards that completion instead of starting another
main-loop iteration. A request-hook crash does not send the prompt.
Hooks let you add optional steps to the Mantle turn path without replacing core
components. Each hook is an async function registered with a point-specific
decorator. Observe hooks (on_*) do not block the turn; if you register several
for the same point, they run in parallel. Modify hooks (modify_*) are awaited
and may return an updated payload.
Layout
Place hook modules at the agent root:hooks.py and nested
hooks/**/*.py files (excluding __init__.py, __pycache__, and *.pyc).
Shared project code under lib/ is importable during hook module loading, the
same way as tool modules.
Registration
Import the decorator you need directly fromrasa.mantle.hooks. Parentheses are
mandatory, even when you only use defaults:
hooks/redact_pii.py
- One modify hook per point. A second registration fails at train time and model load.
- Many observe hooks per point. They run concurrently and unordered. The engine does not await them on the turn path.
- Always
async def. Synchronous handlers are rejected at validation. - Signatures: observe hooks take one typed payload and return
None; modify hooks take one typed payload and return the same type. - Control flow is raised, never returned. Use exceptions for vetoes and model steering (see below).
Insertion points
Incoming-message and outgoing-text hooks apply to text only. Hooks for incoming
audio, ASR partials, and outgoing audio are not available.
Control signals
Modify hooks signal intentional non-happy paths by raising exceptions:HookRejection(reason=...)— veto this action. Allowed only at points listed in the table above.RetryModel(feedback=...)— discard the current model response or tool call and steer the agentic loop with feedback. Allowed only onmodify_model_responseandmodify_tool_call.
RetryModel.feedback is a short note the engine adds to the next model request
in the same turn. It is not saved in the conversation history, never appears as
a user or bot message, and is never sent to the user. It only guides that next
in-turn model call.
Multi-tool batches
There is still one registeredmodify_tool_call for the agent. When
an LLM response contains several tool calls, the engine treats that single
modifier as a preflight phase:
- The engine invokes
on_tool_call/modify_tool_callonce per call in the LLM batch, sequentially in batch order, each time with that call’s own payload — before listen-only routing, the knowledge_searched latch, or any tool in the batch runs. - Only after every preflight invocation succeeds does routing and tool dispatch begin.
- Confirmation and tool-constraint gates still run inside tool dispatch after preflight. A closed gate never reaches the tool body; confirmation resume later uses the (possibly hooked) stored arguments and does not re-run these hooks.
RetryModel, the entire batch is
discarded, no tool in that batch runs, accompanying assistant text is not
emitted, and the feedback steers the next model iteration. This prevents a
retry requested for a later call from following an earlier tool side effect.
Parallel preflight is out of scope for the first tool-call wiring.
Unexpected modify_tool_call crashes or timeouts are fail-closed: the batch
is not dispatched, accompanying text is not emitted, and the user hears
utter_tool_call_hook_error
(overridable in responses.yml). The exception text is not spoken.
After each dispatched call returns, on_tool_result / modify_tool_result
may rewrite the value persisted on tool_executed and
replayed to the next provider request. That point has no control signals;
unexpected crashes are fail-open and keep the original result. Tool-call ids
and batch metadata are preserved.
Tool-call and tool-result hooks run on the LLM tool-call batch path only.
They do not run for other engine-driven tool invocations such as
confirmation resume, settable writes,
memory validation tools, or correction replay.
Knowledge search
on_knowledge_query / modify_knowledge_query run after the empty-query guard
and before the references index is searched. on_knowledge_results /
modify_knowledge_results run after retrieval and before passages are
serialized back to the model. Both modifiers are change-only: they may
rewrite the query or filter/reorder/empty passages, but they do not accept
HookRejection or RetryModel.
Unexpected query-hook crashes or timeouts are fail-closed and skip the search
(the model receives the same safe empty-result note used when retrieval fails).
Unexpected results-hook crashes or timeouts are fail-open and keep the original
passages. Missing-index, empty-query, and retrieval-exception paths are
unchanged when hooks are absent or succeed.
KnowledgeQueryPayload / KnowledgeResultsPayload fields:
These hooks cannot block the turn. Do not raise
HookRejection or
RetryModel here. If you do, Mantle treats that as invalid modify behaviour.
If modify_knowledge_query crashes or times out, Mantle skips the search. If
modify_knowledge_results crashes, times out, or returns an unexpected shape,
Mantle keeps the original passages.
Working example
hooks/rewrite_knowledge_search.py
Rejections and failures
Contract for incoming-message / outgoing-text wiring: Whenmodify_incoming_message rejects text, the original user message remains
the source-of-truth tracker event but is marked discarded. The LLM conversation
projection excludes or safely masks it, and the user receives an
incoming-message fallback. Fail-closed crash/timeout also must not run the
engine on the original payload.
When modify_outgoing_text rejects text, the original response is not emitted.
A discarded bot utterance is stored and the user receives an outgoing-text
fallback. An unexpected outgoing-text hook crash or timeout remains fail-open,
so the original text is stored and sent.
A point’s fail policy applies only to unexpected crashes and timeouts, never
to intentional control signals. closed stops the current turn and fails
gracefully; open logs and continues with the last good payload. Observe hooks
are always non-blocking: they receive independent copies, and their mutations,
errors, timeouts, or control signals have no effect on the turn.
Timeouts
Default timeouts:
Override per hook:
@modify_model_request(timeout_ms=2000).
Payload wall
Every payload carries read-only identity fields:sender_idcall(telephony/session metadata, when present)channel(channel name and front-end metadata)
payload.model_copy(update={...}) to produce modify-hook return values:
Turn lifecycle
on_turn_started and on_turn_finished are observe-only. They run once per
Mantle turn: started when the turn begins, finished when it ends — including
when the turn fails unexpectedly — so you always see a matching pair.
Neither point has a modify lane. Observers are scheduled concurrently and are
not awaited on the turn path. Crashes, timeouts, mutations, and control signals
from observers are logged and discarded — they cannot change tracker or engine
state.
TurnStartedPayload fields:
TurnFinishedPayload fields:
TurnFinishedPayload.status reports the turn outcome:
Working example
hooks/turn_timing.py
Validation
Hook modules are validated duringrasa train (before packaging). Validation
checks:
- import and syntax errors
- unknown hook points
- duplicate modify registrations
- non-async handlers
- payload type annotations and return types
- unsupported control signals for each point
modify_model_request
This wired insertion point runs once per main LLM loop iteration, and on
rephrase, confirm-rephrase, and the post-activate set_fields fill, after
the prompt is trimmed to the token budget and before the provider call. A
crash here does not send the prompt. Empty-completion retries reuse that same hooked request
(messages and allowlisted model_settings) without re-running hooks — the
messages do not change between those attempts, so the hook result stays in
effect for each retry.
ModelRequestPayload fields:
You can change these
model_settings keys: temperature, top_p,
max_tokens, presence_penalty, frequency_penalty, stop, and seed.
Values must match provider-safe types (numbers for the float/int keys; a string
or list of strings for stop). Anything else is ignored — including API keys,
the provider, and the model name. Hooks also cannot change which tools the LLM
is offered.
Message pass-through. Returned messages are forwarded to the provider as
dict copies — Mantle does not validate roles/content or strip extra keys.
An empty list still triggers the LLM call (the failure is a provider error, not
HookExecutionError). Extra fields such as tool_calls or name are left in
place intentionally so builders can shape provider-compatible payloads.
This hook cannot block the turn. Do not raise HookRejection or
RetryModel here. If you do, Mantle logs and ignores them, keeps the last good
request, and still calls the LLM. If the hook crashes or times out, Mantle
skips the LLM call and tells the user with
utter_model_request_hook_error.
Working example
hooks/lower_temperature.py
modify_model_response
Runs once per main LLM loop iteration, and on rephrase, confirm-rephrase, and
the post-activate set_fields fill, after the provider returns and before
any user-visible text emit or tool dispatch. The modify lane may rewrite
response text and tool calls; those rewritten values are what the engine uses.
On the three quiet calls, RetryModel discards that completion. It does not
start another main-loop iteration.
ModelResponsePayload fields:
Control signal: raise
RetryModel(feedback=...) to discard this response
(no text emit, no tool dispatch) and steer another model iteration. The
feedback string is appended as an engine-only system note on the next
request only. It is not stored on the tracker and is never sent to the user.
Re-entry still counts toward the per-turn iteration cap; exhausting that cap
uses the existing max-iterations result.
Do not raise HookRejection here. Once an incoming message has been
accepted, this point improves accuracy/safety by rewriting or steering — it
cannot intentionally suppress the answer. An unsupported rejection is logged
and discarded, and the original response proceeds.
Unexpected modifier crashes, timeouts, or invalid return values are
fail-open: Mantle keeps the original response and continues. RetryModel
is intentional control flow and is never swallowed by fail-open handling.
When a modify_model_response or modify_tool_call hook is registered,
Mantle buffers the full main-LLM completion instead of streaming tokens live.
Those hooks run only after the provider finishes; live streaming would let the
user hear tokens the modifier might still rewrite or discard via RetryModel
(including text dropped with a discarded tool-call batch). Channels that
support streaming still receive the final (possibly hooked) text in one
delivery. Observe-only hooks do not disable streaming.
Changing between text-only and tool-call shapes also updates the
knowledge_searched latch used for the search-then-answer loop — the latch
follows the post-hook tool calls.
Working example
hooks/ground_answers.py
modify_tool_call
Runs as a batch preflight before listen-only routing, the
knowledge_searched latch, and any tool dispatch. For each call in the
batch, Mantle invokes on_tool_call / modify_tool_call once — sequentially
in batch order — with that call’s own payload. Only after every preflight
invocation succeeds does routing and dispatch begin.
Confirmation resume, settable writes, memory validation tools, and correction replay
never reach this path. Confirmation and tool-constraint gates still run later
inside dispatch after preflight.
ToolCallPayload fields:
Engine-owned tool-call
id and provider type are preserved across rewrites;
hooks never receive or change them.
Control signal: raise RetryModel(feedback=...) from any preflight
invocation to discard the entire batch (no tool has run yet, and any
accompanying assistant text is not emitted) and steer the next model
iteration. The feedback string is appended as an engine-only system note on
the next request only.
Do not raise HookRejection here. An unsupported rejection is logged and
discarded, and the original call proceeds through preflight.
Unexpected modifier crashes or timeouts are fail-closed: the batch is not
dispatched, accompanying text is not emitted, and the user hears
utter_tool_call_hook_error.
A registered modify_tool_call hook buffers the main completion the same way
modify_model_response does, so live streaming cannot leak speech that
preflight later discards. Observe-only on_tool_call hooks do not disable
streaming.
Working example
hooks/guard_tool_calls.py
modify_tool_result
Runs after each dispatched tool returns and before the result is persisted
on tool_executed and replayed to the next provider
request. The modify lane may rewrite value, arguments, and is_error.
ToolResultPayload fields:
Tool-call ids and batch metadata stay engine-owned.
This hook has no control signals. Do not raise
HookRejection or
RetryModel here. If you do, Mantle logs and ignores them and keeps the last
good result.
Unexpected modifier crashes, timeouts, or invalid return values are
fail-open: Mantle keeps the original result and continues.
Langfuse tool-call span output uses the post-hook persisted result (same
value stored on the tracker), so redaction in modify_tool_result applies to
traces as well as prompt replay.
Working example
hooks/redact_tool_results.py