Skip to main content
agent.yml is required at the project root. It carries the agent’s identity in an agent: block, plus several top-level tuning sections.
agent.yml
The file has two levels. id, language, and persona are the agent’s identity and live inside the agent: block. rules, prompts, references, skill_retrieval, conversation, tool_timeout, orchestrator, and session_config are tuning sections and sit at the top level as siblings of it.

The agent: block

Required. A missing or non-mapping agent: section fails the load.

persona

Required, together with orchestrator.model_group. A missing or blank persona raises config.agent.missing_persona at load. It sets global tone and identity.

id

When absent or blank, a unique id is generated in the form YYYYMMDD-HHMMSS-<duoname> and written back into agent.yml, with a warning. language is not written back. The en default is applied at parse time only.

rules

A list of strings, rendered into the system prompt as a bullet list. These are global do/don’t guidance applied across every skill: scope limits, tone guards, ordering constraints.
A rule containing : followed by a space parses as a YAML mapping and breaks the load. Quote it, or use a > folded scalar as above.

prompts

Overrides for individual sections of the system prompt. Every key is optional; an empty or absent value falls back to the built-in default.
routing_no_active_skill is the lever for greeting/closing behaviour and for telling the orchestrator when not to route into a skill.

references

Controls how the agent answers from retrieved knowledge. The same model group embeds the corpus at rasa train and loads it at serve time, so keep it declared for as long as the trained model is in use.

chunking

At rasa train, each reference .md file is split with recursive character chunking, then each chunk is embedded. search_knowledge returns those sections, not whole files. Size and overlap are characters, not embedding tokens.
agent.yml
Keep chunk_size well under the embedder’s maximum input. Azure OpenAI embeddings and OpenAI text-embedding-3-small / text-embedding-3-large cap input at 8192 tokens. The default 1000 characters stays under that limit; a much larger character size can still fail train. Changing chunking requires rasa train. Serving does not re-read references/ from disk. See References for when to override the defaults.

skill_retrieval

Optional top-level block. When it is on, each user turn offers only a retrieved subset of independently startable skills in the routing catalogue and the activate tool — not the full list. This is not knowledge search. references / search_knowledge retrieve passages from references/ to answer questions. skill_retrieval chooses which skills the model may start. Omit the block to keep retrieval on with the defaults below. Set active: false to keep today’s full legal catalogue (no index is written). Models trained before the index existed must be retrained, or set active: false and retrain. Silence-timeout turns do not search; they keep always-include, previously started, and the live skill. Unknown keys under skill_retrieval fail the load.
agent.yml
At rasa train, independently startable skills (not engine-managed, not disabled, not hidden, with a routable entry) are indexed. Those same filters apply at runtime, so a retrieved or pinned skill stays out of the catalogue when it is not legal to route. Skills with always_include_in_prompt stay in the catalogue when they are legal, even if search missed them. Skills started in the last 20 utterance lines of the current conversation stay offered the same way. A restart or a new session starts a new window. If search or embedding fails at runtime, the turn fails the same way as an LLM outage. The agent does not fall back to listing every skill. A train with active: false and more than 20 independently startable skills logs a count-only warning.

conversation

before_end lists skill ids that should complete before the conversation ends. The names are rendered into the routing prompt, so the orchestrator knows to pick those skills up before wrapping up. Voice silence is engine-managed: three check-ins, then goodbye and hangup on the next timeout. Any user speech resets that counter. Builders can override the bundled silence responses; the retry count is not configurable. Default wording is on Silence timeout.

tool_timeout

Optional top-level key. Wall-clock timeout in seconds for tool calls.
agent.yml
When a tool exceeds the limit, the runtime cancels the call and returns an error payload to the LLM so it can recover (for example, retry or ask the user). Raise this value for slow backends (databases, long HTTP calls); leave it unset to keep the 10 second default.

session_config

Voice

Voice is configured as a channel in integrations.yml, including the ASR and TTS engines under that channel’s asr: and tts: keys. The runtime detects a voice channel and applies prompts.voice_rules from this file, so what belongs here is the persona and the voice channel rules.

Starting a conversation

The bundled default_session_start skill does not greet. Who speaks first decides whether you should override it. The other bundled skills and their default wording are in Default skills.
  • Agent speaks first — override default_session_start with a greeting and have the channel send /session_start as the first user message.
  • User speaks first — leave the bundled skill silent so the first real utterance can be handled immediately.
rasa shell, the Inspector, and the built-in voice channels already send /session_start. See Start a conversation.

orchestrator

Names the model group that runs the agent, and binds skill precondition names. The group itself — provider, model, and API key — is defined in integrations.yml model_groups.
agent.yml
A missing orchestrator: block, or a blank model_group, fails the load. Changing model_group needs a retrain. Changing the models inside that group in integrations.yml needs a restart. A leftover llm: section in integrations.yml is invalid.

preconditions

Maps each precondition name to how the engine clears it before a skill may start. The name must match a skill’s precondition: field (for example customer_authenticated on Card Replace). Define an entry here for every precondition: used in the project.
agent.yml
The waiting skill stays in routing. The engine parks it as pending, calls the resolver, then promotes or abandons — no resume prompt. If resolve_with is already the active skill (for example the session opener is also the resolver), the engine parks the waiting skill under that run instead of interrupting it and starting a second resolver. See skill.md precondition.

See also

  • Start a conversation: who speaks first, and when to override default_session_start
  • integrations.yml: model groups, channels, MCP servers (per-server tool_timeout), optional tracing and metrics
  • Tools: @tool interface and timeout behaviour
  • memory.yml: project-wide memory
  • skill.md: always_include_in_prompt and other skill frontmatter