How it works
Drop markdown files in areferences/ folder. At rasa train, every .md file
is split into overlapping character chunks, then those chunks are embedded into
a local vector index that ships inside the model. At runtime the agent gets a
built-in search_knowledge tool and calls it when a question needs looking up.
Retrieved hits are the matching sections, not the whole file.
references/ folder exists at train
time, search_knowledge appears. Optional embedder and chunk-size settings
live under references in agent.yml.
Markdown is the indexed format: every **/*.md file under a references/
folder, with empty and whitespace-only files skipped.
Both locations feed one index
Files at the agent root and files underskills/<id>/references/ are indexed
together into a single searchable store, and every search_knowledge call
covers all of it whatever skill is active. The two locations are there to keep
your files organised alongside the skill they belong to.
Write each document so it stands on its own, then. A retrieved chunk arrives
without the folder it came from, so a fees document that opens with “Billing
fees for postpaid plans” is usable in a way that one opening with “The
following fees apply” is not.
A large FAQ can stay in one markdown file. You do not need to split it by hand
before training: index-time chunking does that. Headings and short sections
still help retrieval, because each stored vector is one chunk.
What happens in a turn
- The user asks something.
- The model calls
search_knowledgewith a natural-language query. - Matching sections come back as a tool message, each with its source file.
- The model answers from those sections, or calls
search_knowledgeagain with a refined query.
- After a search, the skill’s own tools and
listenare withheld for the rest of the turn.activate,search_knowledge, andcannot_helpremain, plusresolve_tool_confirmationwhen a confirmation is pending, so a grounded answer neither re-drives the active skill nor ends before presenting the retrieved information. The skill’s tools andlistencome back on the next turn. cannot_helpis withheld until a search has run. When a knowledge base exists, the agent cannot decline a request before retrieval has had a chance to answer it.
Choosing the embedding model
By default the index uses the built-in OpenAI embeddings. To choose another, declare a model group and name it:integrations.yml
agent.yml
rasa train and loaded at serve time with the same
embedder, so keep the named model group in integrations.yml for as long as the
trained model is in use. Renaming or removing it means retraining.
How files are split
Each markdown file is split with recursive character chunking before it is embedded. Size and overlap are measured in characters, not embedding tokens. Omitchunking to use the defaults (chunk_size: 1000, chunk_overlap: 20).
Override them when a document’s sections are much shorter or longer than that:
agent.yml
chunk_size well under the embedder’s maximum input. Azure OpenAI
embeddings and OpenAI text-embedding-3-small / text-embedding-3-large reject
inputs over 8192 tokens. A 1000-character default stays under that; a very
large chunk_size can still fail train.
Changing chunking has no effect until you run rasa train again. Field names,
defaults, and validation are in agent.yml.