← Yang Cao 12 min read 10.09.2026

Eight decisions behind a self-serve data agent

A self-serve data agent whose answers are grounded in sources you can open. What powers it, the eight decisions that made the answers high quality, and what a thousand real questions revealed about the data underneath.

“Coding is an open-ended solution space that rewards the models’ creativity, while documentation and tests provide natural guardrails against hallucination. In contrast, for analytics use cases, there’s often only a single correct answer using a single correct source in which there’s no deterministic way of proving the correctness.” Anthropic, on their own data team’s agent.

That is the whole difficulty in one place: nothing downstream tells a data agent it was wrong.

And yet the loop that runs the self-serve data agent my team and I built in a two-day hackathon, and that I have maintained since, is seven lines: ask the model, run the tool it picks, hand the result back, ask again. Nothing that makes its answers good is in those seven lines. What makes them good is curated tools with rich descriptions, and grounding in sources we control, and the rest of this post is that sentence taken apart.

First, the receipt. The agent lives in Slack. It is now an official product, the first to come out of the engineering team. Since June it has answered 1,057 questions on 64 active days, from 57 people across seven channels, and no onboarding doc exists; people tagged it, it answered, and they told each other. 87% of its answers point at a source you can open, which is the property I care about most and the one the rest of this post is organised around.

By the end you should know three things: what powers it, the eight decisions that made the answers high quality, and what a thousand questions revealed about the data underneath. That should be enough to build one like it yourself, and I would like you to, because the interesting problems turn out not to be in the agent at all.

What powers it

Ask. Run. Hand back. Ask again.

A question arrives in a Slack thread. The model receives it together with ten tools spread across four sources: the official documentation, the application database, the source code, and the data warehouse. It asks for one tool. The code runs it, hands the result back, and the model asks again, 6.8 rounds on average, until it stops asking and whatever it has written is the answer, with its sources attached and a note of what it did not check.

Question a Slack thread Cited answer + what it did not check Model one frontier model A tool runs one of the ten below asks for a tool result handed back ↻ 6.8 rounds on average Official docs the rule as written search_docs get_page App database the records, read-only query_db describe_table The code the rule as implemented grep_code read_file list_files Warehouse history and scale query_warehouse describe_table list_tables nothing decides which box to reach for: the model reads all ten descriptions and picks
Fig. 1: one question in, one cited answer out. Four sources, ten tools, one loop, no router.

That circle is the entire loop, and it is seven lines of code. There is no framework underneath it, no knowledge base, no vector database, and nothing that decides which of the four boxes to reach for. I am not claiming any of this is clever. The loop is free, and everyone gets the same one. The months went into what the loop reads.

Eight decisions

Few questions, answered well

Four sources are wired in: the official documentation, the application database, the source code, the data warehouse. The documentation is the official one, the version written for our customers and kept accurate, not just any document that mentions the topic. Four are left out on purpose: the internal wiki, where a page from two years ago and last week’s decision look the same; free search over Slack; the open web; and anything I cannot vouch for. That is not a judgement on those sources. It is that I cannot control what comes back from them, and the arithmetic of trust is lopsided: one confidently wrong answer costs more than ten questions the agent politely declines to answer. So the wiring stays narrow.

Rich descriptions, not predefined routes

There is no router. The obvious design puts a classifier in front of the tools to decide whether a question is about code or data or documentation, and the obvious design is an if-else in disguise: every new scenario becomes another branch, and it degrades exactly as the scenarios multiply. Instead the model reads all ten tool descriptions and decides for itself where to look. Adding a source means writing one description. Fixing a routing mistake means sharpening a sentence. Descriptions scale in a way a classifier never will. They are also only the visible edge of where the time goes: most of it goes into the data foundation underneath, organising the datasets, modelling the data more logically, and adding metadata and descriptions to the datasets themselves.

A router Rich descriptions Question Classifier if-else in disguise ✗ wrong guess Question Model reads all ten, picks every description in view
Fig. 2: a router is an if-else in disguise. Rich descriptions put every tool in view and let the model pick.

Grounded and reproducible

Every answer has to be checkable by the person reading it, and I mean enforced by code rather than requested in a prompt. Three mechanisms do it. The sources are recorded by code: the loop logs every source it opens, and the model cannot edit that record, so it cannot cite a page it never fetched. If it finishes with no sources, the code says so and stamps the answer as from general knowledge, not verified. And every query is kept verbatim as it ran, behind a button on the answer that replays the chain. The References section you read in Slack is the model writing. The recorded log underneath is the control, and the log is the thing I audit.

Clone the repo. Do not wrap it in a protocol

For code, the agent greps a local clone, and nothing sits between it and the source. A grep is a grep: a regular expression over the files exactly as they are on disk, and a byte-range read of any of them. GitHub’s code search API can do neither; it matches keywords against an index rebuilt after each push, so it is both less precise and a little behind. There is also nobody in the hot path: no rate limit, no auth, no third-party outage arriving halfway through an investigation. And the tool list stays ours. An MCP passthrough once grew a write tool between two deploys, and the agent used it, unprompted, while nothing had changed on our side. If you use Claude Code you already trust this mechanism, because ripgrep over a local checkout is what makes it good at code. The same thing works here.

Guardrails, written as refusals

The agent’s guardrails are its permissions, not a paragraph of English in the prompt. It has three identities of its own: a cloud role with no stored key that rotates hourly, for the warehouse; a read-only database role, for the application database; and a GitHub App whose token expires hourly and can see three repositories and no others. The green arrows in the figure are what each identity may do. The red ones are the product.

Cloud role no stored key, rotates hourly query · read ✗ INSERT, CTAS the data warehouse Database role read-only SELECT ✗ any write the app database, read replica GitHub App token expires hourly clone · read ✗ push three repositories, no others every red arrow is proven by a live check that runs the forbidden thing and records the denial
Fig. 3: three identities of its own, and what each is refused. The red arrows are the product.

Every blocked arrow is proven by a live check that runs the forbidden thing and records the denial.

Safe to deploy, fast to change

I wanted this from day one: a CI strict enough that merging is boring, because boring merges are what let one person deploy continuously and iterate fast without lying awake. One question sits under every check: did it run in the environment the code ships to? The pod runs the image, and a green run on my laptop says nothing about the image. So the checks are organised by where they run.

When What runs Where Proves Auto
Every commit 433 unit and component tests, every source mocked laptop, pre-commit hook the functions behave; a red suite cannot land yes
Every push the same tests, then two assertions inside the built image CI, inside the image the image holds its search program and its grounding files yes
Before a merge live checks: capabilities, fences, identity, grounding laptop for a baseline, the pod for proof every tool reaches its source; no write path to production data; the caller is the service role; the files are in the image no
Every start startup check the pod grounding files present or no start; which sources are reachable yes
After a merge deploy verification laptop, driving the pod merged commit to verified production pod, ending with a real question to the live agent no
Answer quality question bank of real colleague questions with human-verified answers inside the image, read by eye an answer is right, beyond the plumbing working no

Each row is one command, wrapped as a Claude Code skill so nobody has to remember how, and shipping rides the same GitOps path as every other service. What it buys: one person shipping a change to production the same day, repeatedly.

The last column is the honest part. Half the rows still need a person to start them and read the result, there is no automatic scoring of answers, and there is no test environment: the pod in production is the only place the service identities exist, so the deployment proof can only run there. Moving each of those into the pipeline, with a real test environment in front of production, is the work in progress.

Where I want this to go is one step further: an agent that learns from its own corrections, and knows where each lesson belongs. When a colleague corrects an answer, the cause is one of a few things: a badly designed table, a tool that is missing, a tool description that is unclear, or a piece of tribal knowledge nobody wrote down. Each of those has a different home, the data model, the code, a description, the documentation, and remembering the correction is worth little unless the fix lands in the right one. So the memory I want is not a store of past answers; it is the agent proposing the change at the right layer, the same gates running against it, and a human reviewing the intent at the top while the fences hold underneath. The gates exist so that whoever maintains this can be bold. Eventually that does not have to be a person.

Start simple: zero infrastructure

The agent is one pod that dials out to Slack over one websocket, and nothing comes in: no endpoint, no webhook, no address to attack. The thread is the memory, re-read on every mention, so nothing is stored anywhere. That drawing has not changed since the hackathon, and since the move to the cluster there have been zero outage mentions in roughly 400 questions. Nothing to host, nothing to store, nothing to be paged about.

Telemetry from day one

Every answer is recorded: the tools, the rounds, the exact SQL.

Where it looks tool calls, June to September App database in 55% of answers query_db 3,876 describe_table 1,107 Warehouse in 34% of answers query_warehouse 1,737 describe_table 651 list_tables 237 The code in 28% of answers grep_code 1,024 read_file 726 list_files 95 Official docs in 13% of answers search_docs 199 get_page 72 How long it thinks answers by rounds before the model stopped asking 1 5 10 15 20 25 6.8 rounds on average 158 answers needed one round; 7 ran to the last
Fig. 4: two things the record answers without new instrumentation. Where the agent looks, as tool calls grouped by source with the share of answers that touched each; and how long it thinks, as rounds per answer.

Out of that one record come the things I actually wanted to know: what people ask, where the gaps are, what data is missing, which of our own words confuse us, and how to tune the agent’s tools, rounds and cost. The figure is two of those, read straight off the log: the application database is where most questions go, the official documentation is where the fewest do, and most answers finish in a handful of rounds with a long tail that runs to the cap. Build the measuring before you need the measurement. Every number in this post came out of it.

The one line to leave with

Spend your time on the data foundation: the sources you ground in, and what they say about themselves. Not on the loop.

Grounding, precisely, means the answer comes from a live call against a source we own rather than from the model’s memory. Anthropic’s data team ended up in the same place with their own agent.

“For self-service agentic business analytics, the complexity mainly lies in the ambiguity of the data. The central problem comes down to our ability to map a user’s question to specific and up-to-date entities in our data model and know the correct way of working with them. If we can do that, then the resulting execution and SQL becomes trivial.” Anthropic, on their own data team’s agent.

What people actually ask

I read the first 518 of the 1,057 questions by hand and gave each one a theme. Six themes cover most of them.

Theme Questions
Record lookups and exports 123
How a number is calculated 75
Ad-hoc statistics 72
Analysis for external partners 45
Support debugging 44
Live progress tracking 44

Two of every three questions ask the agent to query our databases directly.

The one gap

238 of those 518 questions needed somebody to say which source is the authority.

75 were answered by reading the source code: questions about how a number is calculated, from 20 colleagues in all seven channels, where the rule is written down nowhere except the code itself. Another 171 needed the agent to reverse-engineer a table before it could answer at all. On one request it produced five successive answers from five different tables, each of which looked authoritative, and every correction came from one colleague’s memory.

We do document our tables. What we never wrote down is which table is the authority for a given business question, or what its words mean, and one everyday term turns out to name three different things in our own database.

None of that is the loop. All of it is the data foundation, and that is where the interesting work is.

Comments