The assistant eval suite runs the real agent against a live LLM and scores the outcome. Evals are defined in the codebase, synced automatically into Phoenix as datasets, and every run is recorded there as an experiment with per-case scores, cost, latency, and a full trace. Why this platform: ADR 007.
.env.docker-dev: COMPOSE_PROFILES=optional,ai,ai-evals, plus a
provider API key (e.g. GROQ_API_KEY) — see
AI assistant tracing for the
Phoenix side.just dc-dev up -d — the assistant-eval-runner service migrates its own
baserow_evals database, syncs the datasets into Phoenix, and serves
http://localhost:8090.running 7/38; click it to expand the
run’s log tail. Each queued or running row has its own Stop, and
Stop all halts everything at once.The status cell counts finished case-repetitions against the total, which is known before the first case starts. Clicking any row — running or finished — expands the last 200 log lines captured from that run: one line per case with its score and duration, plus any warnings and tracebacks. Only the worker thread’s lines are captured, so page requests never pollute a run’s log.
The page polls every 2s while anything is queued or running and backs off to 30s once everything settles, repainting only the rows whose state actually changed. A finished run’s log is fetched once and then left alone.
The buffer is in memory and deliberately not persisted: a .py edit restarts
the runner, which rewrites in-flight runs to failed and drops their logs.
Raise or lower the captured level with BASEROW_EVAL_RUNNER_LOG_LEVEL
(default INFO). Only Baserow’s own loguru output is captured — library
chatter (httpx, pydantic-ai retries) still goes to
just dc-dev logs -f assistant-eval-runner.
Every case has a wall-clock budget — BASEROW_EVAL_CASE_TIMEOUT, default
120s. For scale: across the committed baseline’s 111 runs the slowest case
takes 16.4s and the median 5.8s, so the budget only ever fires on a genuine
hang. Without it a single stuck case blocks the one worker indefinitely, and
the per-request timeouts don’t bound it: max_iters requests times the
per-request timeout, plus retries, runs into several minutes.
A timed-out case cancels the in-flight provider call and signals its tools to
stop. Agent cleanup and tool operations already running share up to five
additional seconds to finish, before prompts are restored or another case
starts. Once cleanup finishes, the case is recorded as a failed
completed_within_timeout check with score 0, the run continues, and the judge
does not grade its empty answer.
If cleanup is still running after that period, the experiment fails
and the runner refuses further cases. In-flight threads cannot safely be
stopped, so prompt overrides remain in place. Restart the eval runner before
running another experiment (just dc-dev restart assistant-eval-runner).
Stop is cooperative and lands at the next case boundary, because the
worker sits inside a blocking LLM call that Python cannot interrupt. Queued
runs stop immediately; a running one finishes its current case first and ends
as stopped, keeping the cases it already logged to Phoenix — its status
reads stopping… in between. Stop one dataset from its row, or use Stop
all when an error is going to sink every remaining case anyway.
A selection spanning several datasets fans out to one experiment per dataset,
and the Results tab groups experiments by name. Leaving Experiment name
blank generates one shared run-<timestamp>-<id> name for the whole fan-out,
so the Results tab compares every dataset against the baseline in one view.
Type a name instead to group runs yourself — reusing a name across separate
submissions merges them into one group.
Experiments created before this grouping existed each carry their own Phoenix-generated name, so they stay ungrouped.
Every experiment is stamped with its model, the resolved orchestrator
model_settings (temperature, reasoning effort, max tokens), the judge model,
prompt hashes, and git branch/commit — so a score is traceable to the
configuration that produced it without writing anything down. The Notes
field adds free text for whatever that does not cover; both the settings and
the note show under the model in the Results tab.
Warning: the runner hot-reloads on any mounted
.pychange (including a lint/format pass), which kills queued and running experiments — don’t edit backend Python while a run is in flight.
The CLI uses your host env (.env.local): the database it points at and
BASEROW_ASSISTANT_PHOENIX_URL for Phoenix.
Warning: CLI runs create real scenario data (users, workspaces, apps) in whatever database
DATABASE_NAMEpoints at, with no teardown. The runner service uses its own disposablebaserow_evalsdatabase, so prefer it for bulk runs, or pointDATABASE_NAMEat a disposable database first.
# Sync the datasets defined in code into Phoenix (idempotent)
just b eval-sync
# Run a whole dataset
just b eval-run --dataset kuma-core
# Run selected cases (must all belong to the same dataset)
just b eval-run --case database/creates-simple-table --case database/creates-view-kanban
# Compare models / measure flakiness
just b eval-run --dataset kuma-builder --model groq:openai/gpt-oss-20b --runs 3 --name builder-20b
--runs N repeats every case N times in one experiment — the score spread
across repetitions is the flake signal. Compare experiments (models, prompt
changes, repeat runs) side by side in the Phoenix dataset view. Every
experiment is auto-stamped with its model, git branch and commit, and prompt
hashes, so branch/model comparisons are filterable in Phoenix.
| Dataset | Cases | Covers |
|---|---|---|
kuma-core |
3 | creating/listing databases and automations |
kuma-database |
21 | tables, fields, views, filters, rows |
kuma-builder |
16 | pages, elements, data sources, themes, user sources |
kuma-automation |
7 | workflows, triggers, nodes |
kuma-docs |
64 | docs Q&A via search_user_docs, incl. cannot-do guardrail cases |
kuma-docs needs the embeddings service (ai profile) and a knowledge base
synced into baserow_evals, independently of the main development database.
Set BASEROW_EMBEDDINGS_API_URL=http://embeddings:80 in .env.docker-dev,
recreate the runner to pick it up, then seed its database using the existing
command:
just dc-dev up -d embeddings assistant-eval-runner
just dc-dev exec assistant-eval-runner just b manage sync_knowledge_base
The command indexes the repository’s enterprise/backend/website_export.csv.
Wait for it to finish before running docs cases. If it reports unavailable
embeddings or pgvector, fix that prerequisite and run it again. When the knowledge
base is unavailable, docs cases are skipped: a finished executor does not mean
the evals passed. Check the run log and per-case Phoenix results.
Cases live in
enterprise/backend/src/baserow_enterprise/assistant/evals/datasets/ and are
picked up automatically (registration on import, synced at runner startup or
via just b eval-sync). A case is three parts — a scenario (the Baserow state
the agent starts from), a prompt, and checks:
from baserow.test_utils.fixtures import Fixtures
from baserow_enterprise.assistant.evals.harness import tool_called
from baserow_enterprise.assistant.evals.registry import register_case, register_scenario
from baserow_enterprise.assistant.evals.scenarios import build_database_ui_context
from baserow_enterprise.assistant.evals.types import (
CheckResult,
EvalCase,
EvalRunOutput,
EvalScenario,
)
@register_scenario("database-my-scenario")
def _my_scenario(fx: Fixtures) -> EvalScenario:
user = fx.create_user()
workspace = fx.create_workspace(user=user)
database = fx.create_database_application(workspace=workspace, name="Sales")
return EvalScenario(
user=user,
workspace=workspace,
ui_context=build_database_ui_context(user, workspace, database),
refs={"database": database},
)
def _my_checks(
case: EvalCase, scenario: EvalScenario, output: EvalRunOutput
) -> list[CheckResult]:
return [
CheckResult("called list_tables", tool_called(output, "list_tables") >= 1),
CheckResult(
"answer mentions Sales",
"sales" in output.answer.lower(),
hint=output.answer[:200],
),
]
register_case(
EvalCase(
id="database/my-new-case", # append-only, never rename
dataset="kuma-database",
prompt="Which tables are in the Sales database?",
scenario="database-my-scenario",
checks=_my_checks,
max_iters=10,
)
)
Rules that keep results comparable over time:
scenario.refs are the norm. The harness automatically
prepends a tool_errors_within_budget check (EvalCase.max_tool_errors,
default 0).scenario.pre_state.passed_checks / total_checks; it passes only when every
check passes.You don’t need a PR to add a case: from the dataset editor, or from a trace
span’s “Add Example to Dataset”, add an example directly in Phoenix. just b eval-sync preserves it — it no longer wipes examples that aren’t in the
codebase, only code-owned ones (identified by a case_id in their metadata)
are replaced wholesale.
UI-added examples are runnable: they appear on the runner page under
“Added in the Phoenix UI” in their dataset’s tab (reload the page after
adding one), and they run with the rest of the dataset too. The example’s
input needs a {"prompt": "..."} (or "question") and everything else is
optional metadata:
Add the example to kuma-docs. It runs against the standard docs scenario
with the standard checks (search_user_docs called, at least one source),
and the LLM judge scores answer_quality. Optional fields:
output → {"reference_answer": "..."} — the ideal answer the judge
grades against.expected_keywords — list of strings; adds an
“answer mentions one of” check and informs the judge.Add the example to the matching dataset (kuma-database, kuma-builder,
kuma-automation, kuma-core) and declare what to exercise in its metadata:
scenario — name of a registered starting state (the
register_scenario("...") ids in evals/datasets/*.py and
evals/scenarios.py). Defaults to empty-workspace, a bare workspace —
enough for “create a table called X”-style prompts. Pick a richer scenario
when the prompt references existing objects; its object names must match
what the prompt mentions.expected_tools — list of tool names; each adds a “called <tool>” check.answer_contains — list of strings; each adds a case-insensitive
“answer contains” check.mode (agent mode, defaults to the dataset’s usual one), max_iters,
max_tool_errors — same meaning as on a code case.Checks that assert on database state (rows really created, field types correct) can’t be expressed in metadata — promote the example to code for those. The tool-error budget check always runs, and the full trace is linked from every run, so even a check-less example is useful for experimenting.
To promote a UI-added example to code:
just b eval-export --dataset kuma-docs
This prints a ready-to-paste _register_docs_case(...) snippet per UI-added
kuma-docs example (question from the example, keywords/source patterns from
its metadata if the UI author set them, else TODO placeholders to fill in;
a reference_answer kwarg is included too if the example’s output carries
one). Paste it into datasets/docs.py, pick a real id and keywords, then
just b eval-sync — the UI copy is dropped automatically because its prompt
now matches the code case (matched by exact prompt text, so no duplicate).
Non-kuma-docs datasets have no registration helper to generate from, so
eval-export prints a commented JSON block instead; write the scenario and
checks by hand.
A kuma-docs case can carry an ideal “reference answer” that the LLM judge
grades Kuma’s answer against (see below). Add one in code with
_register_docs_case(..., reference_answer="..."), or curate it directly on
a synced example in the Phoenix UI by editing its output field to
{"reference_answer": "..."} — that’s a normal, versioned edit to the
example, so it survives just b eval-sync: a code case with no
reference_answer never overwrites a live one, it only adopts it, and a
code-set reference_answer always wins over whatever is live.
EVAL_MODELS in
enterprise/backend/src/baserow_enterprise/assistant/evals/models.py is the
candidate list — extend it there. Any pydantic-ai provider:model string
works via --model, or in the UI by picking Custom… and typing it. The
model applies to the whole agent, sub-agents included.
Eval runs use an explicit model profile, so the selected model takes precedence over workspace or instance provider settings without changing them. Model clients are closed when a run finishes or times out.
Per-model overrides live in _MODEL_PROFILES in
enterprise/backend/src/baserow_enterprise/assistant/model_profiles.py, keyed
by exact model name.
The dropdown is not EVAL_MODELS itself but available_models(), which keeps
only the entries whose api_key_env variable is set in the runner’s
environment. A model with no key is silently absent rather than listed and
broken, so adding one to EVAL_MODELS is not enough to make it appear.
docker-compose.dev.yml forwards GROQ_API_KEY, OPENAI_API_KEY,
ANTHROPIC_API_KEY, GEMINI_API_KEY, GOOGLE_API_KEY, and OPENROUTER_API_KEY to
assistant-eval-runner from .env.docker-dev. Set the variable the model’s
entry names, then restart the service.
Gemini is the one asymmetric case: pydantic-ai authenticates with either
GOOGLE_API_KEY or GEMINI_API_KEY, but available_models() only checks
GOOGLE_API_KEY. With only GEMINI_API_KEY set the Gemini entries stay
hidden even though a run would have worked — use GOOGLE_API_KEY, or reach
the model through Custom…, which skips the key check entirely and so
fails at request time instead of hiding.
Kuma’s load-bearing prompts (the main system prompt and each sub-agent’s
instructions — see SYNCED_PROMPTS in
enterprise/backend/src/baserow_enterprise/assistant/evals/prompt_sync.py)
are synced to Phoenix’s Prompts tab as versioned prompts on every
eval-sync. Every experiment’s prompts metadata records the content hash of
each prompt as it ran, so prompt-version comparisons are filterable in
Phoenix.
To experiment with a prompt change without touching code (Phoenix has no plain edit box — editing goes through its playground):
--override-prompt <name>
(repeatable). The list sorts the active tab’s likely-relevant prompts
first, but any prompt can be overridden — one the selected cases never
exercise is just a no-op.prompt_overrides list naming what was overridden.Editing the constant in code directly still works too (the runner hot-reloads .py changes).
Every experiment run links to its trace (agent → LLM calls → tool calls, with
token counts and cost). Failed checks appear in the experiment’s checklist
evaluator explanation with their hints. How to compare against the committed
baseline, classify outcomes (improvement / regression / gap / flake), and
diagnose failures through traces:
evaluating results.
kuma-docs runs get a third score, answer_quality, from an LLM judge (the
judge prompt lives in evals/judge.py) that grades the answer’s correctness,
helpfulness, and groundedness against the sources the assistant cited. When
the case (or its synced example) carries a reference_answer, the judge is
also given it and told to weigh factual agreement with it heavily — it’s the
ideal answer, not the only acceptable phrasing, so wording differences alone
don’t cost points. The judge model is BASEROW_EVAL_JUDGE_MODEL, defaulting
to groq:openai/gpt-oss-120b, and is stamped into every experiment’s
metadata. A judge failure (LLM error, missing case, …) records no
answer_quality score rather than a 0, so it doesn’t skew aggregates.