AI assistant routing and action memory

Kuma uses one agent with database, application, automation, and explanation modes. The implementation lives in baserow_enterprise.assistant. Modes control tool discovery and schema size; the domain services still enforce the acting user’s permissions and workspace boundaries.

Tool discovery

AssistantToolRegistry first filters tool groups through can_use. The permitted tools share one catalog and a routing map in tools/routing.py. Tools active in the current mode expose their full schemas. Other permitted tools are deferred: pydantic-ai refuses a call to one until the model’s tool search reveals it. search_user_docs is active in every mode, so docs questions skip that search.

Calling a revealed deferred tool changes the mode. If its arguments already match the full JSON schema, it follows the normal validated execution path immediately. Incomplete calls return changed: false with instructions to reissue the call using the full schema, without consuming the tool-error retry budget. Those pending calls remain visible until reissued, and routing-only results are excluded from verified action memory. Tool execution is sequential so later calls see earlier results and mode changes. Dynamic row tools stay available in every mode once loaded and are refreshed when their table schema changes. Loaded row-tool creation totals are per table and per user turn, and survive these refreshes. AssistantDeps.resource_changes owns this runtime tracking separately from persisted action memory.

When adding a tool, register it in its domain’s tool functions and ensure the routing map assigns it an owner. A mode switch never grants permissions or makes an unavailable tool group discoverable.

Action memory and completion evidence

Chat history is compacted to user prompts and final answers, retaining at most 20 messages. action_memory.py also stores recent mutation outcomes in versioned message metadata. This ledger retains at most 12 outcomes and 4,000 serialized characters. Large values are truncated, and row writes keep row IDs and counts instead of cell values, while request fingerprints use the full arguments to distinguish similar requests. The ledger reaches the model as data it must not follow as instructions. An individually oversized outcome can evict earlier outcomes and retain only status flags, losing resource IDs. Improving that compaction is follow-up work; tools can rediscover resources from their current state.

The ledger provides prior resource IDs and partial or failed outcomes to the model. It is bounded context, not an audit log or an idempotency guarantee. Old chats without this metadata remain readable. Tools must still inspect current resource state and enforce permissions before acting.

Completion validation uses only live tool results after the latest user prompt. A prior success, a reused resource, or an explicit no-op cannot by itself justify claiming a new change. A partial result can support an answer that acknowledges unfinished work. Output-validation retries stay within the same user turn. When they run out, Kuma says its answer is unfinished and keeps the turn’s tool results.

The answer checks are English text heuristics. They catch common unsupported success claims, printed tool calls, false mode limitations, and unnecessary handoffs. They do not verify every named resource, interpret all languages, or prove that every part of a request succeeded. Persisted-state checks and trace review remain necessary when evaluating task completion. Ambiguous descriptions fail open: a historical passive sentence or a participle such as “Updated cells highlight briefly” is not enough to assert a new change.

Reuse and clarification

Creation tools reconcile exact-name matches within the authorized parent scope. They return existing resource IDs and report conflicting definitions or requested settings that remain unapplied. Reuse does not silently overwrite an existing resource. It also does not prevent duplicates across concurrent agent runs. When requested creation differs from an existing table, application theme, or workflow, Kuma reports the difference and asks before modifying that existing resource unless the user has already authorized those changes. Verified no-op reuse and explicitly authorized continuation do not need another confirmation.

Product questions use documentation without requiring example tables or fields to exist. Inspection requests use read tools. For changes, missing user-owned data must be looked up before asking for clarification; displaying existing data does not authorize creating replacement tables or sample records. When no table holds the records an app should show, Kuma creates nothing and asks whether to create a table with sample data, agree on its fields first, or use data the user points to. An explicit example, demo, or sample app request authorizes its sample data. New build requests with enough context use reasonable defaults for a useful first version. Compact table schemas identify omitted fields and direct the agent to request the full schema before treating a field as missing. Page discovery distinguishes application IDs from page IDs; setup_page populates a page created separately by create_pages. ask_user records a question for the final answer; it does not persist a separate workflow or introduce a new frontend protocol. The next reply resumes through the ordinary chat history.

Builder formulas and documentation retrieval

Implicit formulas inside a collection use its current record. Explicit formulas can still intentionally access an absolute row or another data source. Builder updates reject unsupported properties before applying changes, retain existing formula values when generation fails, and report any separately committed updates as partial. Buttons navigate through click actions; links store their destination directly and can use button styling. Switching a link to a generated custom URL or an image to a generated URL source is saved together with the validated formula. Failed generation preserves the previous active destination or image source; independent changes can still apply.

Knowledge-base synchronization indexes overlapping passages, including document titles, instead of embedding whole pages that exceed the embedding model’s input limit. Passage metadata triggers reindexing of legacy or incomplete indexes. Retrieval combines semantic and lexical matches and limits passages per document. Answers must cite retrieved sources and distinguish supported information from documentation gaps; missing evidence does not establish feature availability. If synthesis returns no valid citation, it reconsiders the same passages once for supported partial information. This fallback keeps the citation checks and can still return no evidence.

Groq GPT-OSS history keeps prior reasoning in the API’s native reasoning field. Pydantic AI otherwise embeds it in visible <think> text, which can trigger malformed final/tool responses. The adapter preserves text, tool calls, and stored history; other providers and models retain their SDK mapping. Provider failures and malformed responses still count against the same eval limits. Groq GPT-OSS 120B uses high reasoning effort for documentation synthesis. This can increase latency and token usage. The documentation role uses the same output-token limit as subagents and keeps the provider’s timeout, as before this change, instead of imposing the action subagent’s 20-second deadline. A docs-specific deadline requires latency measurement in a follow-up. The synthesis fallback is bounded to two calls; request and error limits remain in effect. Orchestration and utility roles keep their settings. Docs synthesis on Groq GPT-OSS 20B and 120B uses strict native JSON output, avoiding invented output-tool names. Typed result validation and retrieved-source checks still apply. Other models retain their existing output protocol.

Model-specific output protocols, settings overrides, and legacy model adapters are registered together in model_profiles.py. Tools request settings and an output type from the resolved profile using their role, without inspecting model names. Unregistered models keep the role defaults and SDK output protocol.

Verification

Run the assistant and prompt tests with:

just b test ../enterprise/backend/tests/baserow_enterprise_tests/assistant/ ../premium/backend/tests/baserow_premium_tests/prompts/test_prompt_assets.py -n=auto -q

The routing tests exercise discovery, immediate execution of complete calls, deferred execution of incomplete calls, and unavailable groups. History and answer-validation tests cover compaction, eviction, no-ops, failed mutations, and current-turn evidence. Domain tool tests verify saved state and reconciliation. For live model runs, use the eval platform with a disposable database and record the source revision, model settings, case population, and failures.