Status: accepted (2026-08-21). Tracing, the eval framework (111 cases in 5
datasets), and the runner are shipped (2026-08-24); the pytest harness is
retired. Judge evaluators are shipped for kuma-docs (answer_quality);
other datasets don’t have a judge yet.
The AI assistant has ~38 LLM evals that exist only as opt-in pytest tests
(-m eval) run on individual laptops. There is no shared place to run them,
compare models and providers, inspect traces, or track quality, cost, and
latency over time. We need one standard, team-wide way to evaluate the
assistant.
Self-host Arize Phoenix as the eval and LLM-observability platform, plus a small eval runner service we own:
BASEROW_ASSISTANT_PHOENIX_URL is set.
See AI assistant tracing.A structural fact drove the choice: no platform executes your Python agent server-side. “Run evals from a UI” therefore requires a self-owned runner on every platform — which removes the main advantage of the heavier candidates and makes footprint, licensing, and integration quality decisive.
| Candidate | Verdict |
|---|---|
| Phoenix | 1 container, Postgres-backed; zero feature gates; free auth/RBAC/OAuth2; first-party pydantic-ai instrumentation; datasets/experiments/playground/cost tracking included. |
| Langfuse v4 | Full feature set, MIT, native remote-run webhook — but 6 required containers (ClickHouse, Redis, MinIO, …) with no lighter profile. |
| Opik | No authentication at all in self-hosted OSS, ~9 containers, no UI batch runs of a real agent. |
| LangSmith, Braintrust, W&B Weave | Self-hosting is enterprise-contract only. |
| promptfoo, MLflow, agenta, Laminar | Wrong shape (YAML evals, view-only UI, weak velocity, or gated alerts). |
phoenix database in the dev
stack’s existing Postgres (created by the one-shot ai-evals-db-init
service), a dedicated Postgres for the shared team instance. No storage
drift between dev and team.ai-evals, separate from ai
(assistant prerequisites), so using the assistant never requires Phoenix.BASEROW_ASSISTANT_PHOENIX_URL
empty disables the export entirely; PostHog LLM analytics stays the
production path (dual export is a config-only change — see
AI assistant tracing).BASEROW_ASSISTANT_PHOENIX_API_KEY.openinference-instrumentation-pydantic-ai,
later arize-phoenix-client); the export degrades with a logged warning if
they are missing, so production images are unaffected.openinference-instrumentation-pydantic-ai, and
arize-phoenix-client versions are one upgrade unit; pydantic-ai
upgrades can silently degrade trace quality (the span rewriting is
untyped), so re-verify a trace after upgrading either side.