agent.yml is required at the project root. It carries the agent’s identity in
an agent: block, plus several top-level tuning sections.
agent.yml
id, language, and persona are the agent’s
identity and live inside the agent: block. rules, prompts,
references, skill_retrieval, conversation, tool_timeout, orchestrator,
and session_config are tuning sections and sit at the top level as siblings
of it.
The agent: block
Required. A missing or non-mapping agent: section fails the load.
persona
Required, together with orchestrator.model_group. A missing or blank persona
raises config.agent.missing_persona at load.
It sets global tone and identity.
id
When absent or blank, a unique id is generated in the form
YYYYMMDD-HHMMSS-<duoname> and written back into agent.yml, with a warning.
language is not written back. The en default is applied at parse time only.
rules
A list of strings, rendered into the system prompt as a bullet list. These are
global do/don’t guidance applied across every skill: scope limits, tone guards,
ordering constraints.
A rule containing
: followed by a space parses as a YAML mapping and breaks
the load. Quote it, or use a > folded scalar as above.prompts
Overrides for individual sections of the system prompt. Every key is optional;
an empty or absent value falls back to the built-in default.
routing_no_active_skill is the lever for greeting/closing behaviour and for
telling the orchestrator when not to route into a skill.
references
Controls how the agent answers from retrieved knowledge.
The same model group embeds the corpus at
rasa train and loads it at serve
time, so keep it declared for as long as the trained model is in use.
chunking
At rasa train, each reference .md file is split with recursive character
chunking, then each chunk is embedded. search_knowledge returns those
sections, not whole files. Size and overlap are characters, not embedding
tokens.
agent.yml
chunk_size well under the embedder’s maximum input. Azure OpenAI
embeddings and OpenAI text-embedding-3-small / text-embedding-3-large cap
input at 8192 tokens. The default 1000 characters stays under that limit; a
much larger character size can still fail train.
Changing chunking requires rasa train. Serving does not re-read references/
from disk. See References for when
to override the defaults.
skill_retrieval
Optional top-level block. When it is on, each user turn offers only a retrieved
subset of independently startable skills in the routing catalogue and the
activate tool — not the full list.
This is not knowledge search. references / search_knowledge retrieve
passages from references/ to answer questions. skill_retrieval chooses
which skills the model may start.
Omit the block to keep retrieval on with the defaults below. Set
active: false to keep today’s full legal catalogue (no index is written).
Models trained before the index existed must be retrained, or set
active: false and retrain. Silence-timeout turns do not search; they keep
always-include, previously started, and the live skill.
Unknown keys under
skill_retrieval fail the load.
agent.yml
rasa train, independently startable skills (not engine-managed, not
disabled, not hidden, with a routable entry) are indexed. Those same filters
apply at runtime, so a retrieved or pinned skill stays out of the catalogue
when it is not legal to route.
Skills with always_include_in_prompt
stay in the catalogue when they are legal, even if search missed them. Skills
started in the last 20 utterance lines of the current conversation stay offered the
same way. A restart or a new session starts a new window.
If search or embedding fails at runtime, the turn fails the same way as an LLM
outage. The agent does not fall back to listing every skill.
A train with active: false and more than 20 independently startable skills
logs a count-only warning.
conversation
before_end lists skill ids that should complete before the conversation ends.
The names are rendered into the routing prompt, so the orchestrator knows to
pick those skills up before wrapping up.
Voice silence is engine-managed: three check-ins, then goodbye and hangup on the
next timeout. Any user speech resets that counter. Builders can override the
bundled silence responses; the retry count is not configurable. Default wording
is on Silence timeout.
tool_timeout
Optional top-level key. Wall-clock timeout in seconds for tool calls.
agent.yml
10 second default.
session_config
Voice
Voice is configured as a channel inintegrations.yml, including the ASR and TTS
engines under that channel’s asr: and tts: keys. The runtime detects a voice
channel and applies prompts.voice_rules from this file, so what belongs here
is the persona and the voice channel rules.
Starting a conversation
The bundleddefault_session_start skill does not greet. Who speaks first
decides whether you should override it. The other bundled skills and their
default wording are in Default skills.
- Agent speaks first — override
default_session_startwith a greeting and have the channel send/session_startas the first user message. - User speaks first — leave the bundled skill silent so the first real utterance can be handled immediately.
rasa shell, the Inspector, and the built-in voice channels already send
/session_start. See
Start a conversation.
orchestrator
Names the model group that runs the agent, and binds skill precondition
names. The group itself — provider, model, and API key — is defined in
integrations.yml model_groups.
agent.yml
A missing
orchestrator: block, or a blank model_group, fails the load.
Changing model_group needs a retrain. Changing the models inside that group
in integrations.yml needs a restart. A leftover llm: section in
integrations.yml is invalid.
preconditions
Maps each precondition name to how the engine clears it before a skill
may start. The name must match a skill’s precondition: field (for example
customer_authenticated on Card Replace). Define an entry here for every
precondition: used in the project.
agent.yml
The waiting skill stays in routing. The engine parks it as pending, calls the
resolver, then promotes or abandons — no resume prompt. If
resolve_with is
already the active skill (for example the session opener is also the resolver),
the engine parks the waiting skill under that run instead of interrupting it and
starting a second resolver. See
skill.md precondition.
See also
- Start a conversation: who speaks
first, and when to override
default_session_start integrations.yml: model groups, channels, MCP servers (per-servertool_timeout), optional tracing and metrics- Tools:
@toolinterface and timeout behaviour memory.yml: project-wide memoryskill.md:always_include_in_promptand other skill frontmatter