Kuma uses one agent with database, application, automation, and explanation modes.
The implementation lives in baserow_enterprise.assistant. Modes control tool
discovery and schema size; the domain services still enforce the acting user’s
permissions and workspace boundaries.
AssistantToolRegistry first filters tool groups through can_use. The permitted
tools share one catalog and a routing map in tools/routing.py. Tools active in
the current mode expose their full schemas. Other permitted tools are deferred:
pydantic-ai refuses a call to one until the model’s tool search reveals it.
search_user_docs is active in every mode, so docs questions skip that search.
Calling a revealed deferred tool changes the mode. If its arguments already match the
full JSON schema, it follows the normal validated execution path immediately.
Incomplete calls return changed: false with instructions to reissue the call
using the full schema, without consuming the tool-error retry budget. Those
pending calls remain visible until reissued, and routing-only results are
excluded from verified action memory. Tool execution is sequential so later calls
see earlier results and mode changes. Dynamic row tools stay available in every
mode once loaded and are refreshed when their table schema changes.
Loaded row-tool creation totals are per table and per user turn, and survive these
refreshes. AssistantDeps.resource_changes owns this runtime tracking separately
from persisted action memory.
When adding a tool, register it in its domain’s tool functions and ensure the routing map assigns it an owner. A mode switch never grants permissions or makes an unavailable tool group discoverable.
Chat history is compacted to user prompts and final answers, retaining at most
20 messages. action_memory.py also stores recent mutation outcomes in versioned
message metadata. This ledger retains at most 12 outcomes and 4,000 serialized
characters. Large values are truncated, and row writes keep row IDs and counts
instead of cell values, while request fingerprints use the full arguments to
distinguish similar requests. The ledger reaches the model as data it must not
follow as instructions.
An individually oversized outcome can evict earlier outcomes and retain only
status flags, losing resource IDs. Improving that compaction is follow-up work;
tools can rediscover resources from their current state.
The ledger provides prior resource IDs and partial or failed outcomes to the model. It is bounded context, not an audit log or an idempotency guarantee. Old chats without this metadata remain readable. Tools must still inspect current resource state and enforce permissions before acting.
Completion validation uses only live tool results after the latest user prompt. A prior success, a reused resource, or an explicit no-op cannot by itself justify claiming a new change. A partial result can support an answer that acknowledges unfinished work. Output-validation retries stay within the same user turn. When they run out, Kuma says its answer is unfinished and keeps the turn’s tool results.
The answer checks are English text heuristics. They catch common unsupported success claims, printed tool calls, false mode limitations, and unnecessary handoffs. They do not verify every named resource, interpret all languages, or prove that every part of a request succeeded. Persisted-state checks and trace review remain necessary when evaluating task completion. Ambiguous descriptions fail open: a historical passive sentence or a participle such as “Updated cells highlight briefly” is not enough to assert a new change.
Creation tools reconcile exact-name matches within the authorized parent scope. They return existing resource IDs and report conflicting definitions or requested settings that remain unapplied. Reuse does not silently overwrite an existing resource. It also does not prevent duplicates across concurrent agent runs. When requested creation differs from an existing table, application theme, or workflow, Kuma reports the difference and asks before modifying that existing resource unless the user has already authorized those changes. Verified no-op reuse and explicitly authorized continuation do not need another confirmation.
Product questions use documentation without requiring example tables or fields to
exist. Inspection requests use read tools. For changes, missing user-owned data
must be looked up before asking for clarification; displaying existing data does
not authorize creating replacement tables or sample records. When no table holds
the records an app should show, Kuma creates nothing and asks whether to create a
table with sample data, agree on its fields first, or use data the user points to.
An explicit example, demo, or sample app request authorizes its sample data. New build requests
with enough context use reasonable defaults for a useful first version.
Compact table schemas identify omitted fields and direct the agent to request the
full schema before treating a field as missing. Page discovery distinguishes
application IDs from page IDs; setup_page populates a page created separately by
create_pages.
ask_user records a question for the final answer; it does not persist a separate
workflow or introduce a new frontend protocol. The next reply resumes through the
ordinary chat history.
Implicit formulas inside a collection use its current record. Explicit formulas can still intentionally access an absolute row or another data source. Builder updates reject unsupported properties before applying changes, retain existing formula values when generation fails, and report any separately committed updates as partial. Buttons navigate through click actions; links store their destination directly and can use button styling. Switching a link to a generated custom URL or an image to a generated URL source is saved together with the validated formula. Failed generation preserves the previous active destination or image source; independent changes can still apply.
Knowledge-base synchronization indexes overlapping passages, including document titles, instead of embedding whole pages that exceed the embedding model’s input limit. Passage metadata triggers reindexing of legacy or incomplete indexes. Retrieval combines semantic and lexical matches and limits passages per document. Answers must cite retrieved sources and distinguish supported information from documentation gaps; missing evidence does not establish feature availability. If synthesis returns no valid citation, it reconsiders the same passages once for supported partial information. This fallback keeps the citation checks and can still return no evidence.
Groq GPT-OSS history keeps prior reasoning in the API’s native reasoning field.
Pydantic AI otherwise embeds it in visible <think> text, which can trigger
malformed final/tool responses. The adapter preserves text, tool calls, and stored
history; other providers and models retain their SDK mapping. Provider failures
and malformed responses still count against the same eval limits.
Groq GPT-OSS 120B uses high reasoning effort for documentation synthesis.
This can increase latency and token usage. The documentation role uses the same
output-token limit as subagents and keeps the provider’s timeout, as before this
change, instead of imposing the action subagent’s 20-second deadline. A
docs-specific deadline requires latency
measurement in a follow-up. The synthesis fallback is bounded to two calls;
request and error limits remain in effect. Orchestration and utility roles keep
their settings.
Docs synthesis on Groq GPT-OSS 20B and 120B uses strict native JSON output, avoiding
invented output-tool names. Typed result validation and retrieved-source checks
still apply. Other models retain their existing output protocol.
Model-specific output protocols, settings overrides, and legacy model adapters
are registered together in model_profiles.py. Tools request settings and an
output type from the resolved profile using their role, without inspecting model
names. Unregistered models keep the role defaults and SDK output protocol.
Run the assistant and prompt tests with:
just b test ../enterprise/backend/tests/baserow_enterprise_tests/assistant/ ../premium/backend/tests/baserow_premium_tests/prompts/test_prompt_assets.py -n=auto -q
The routing tests exercise discovery, immediate execution of complete calls, deferred execution of incomplete calls, and unavailable groups. History and answer-validation tests cover compaction, eviction, no-ops, failed mutations, and current-turn evidence. Domain tool tests verify saved state and reconciliation. For live model runs, use the eval platform with a disposable database and record the source revision, model settings, case population, and failures.