case study 05

AI agents · evals Concept · live demo
DELTA
AI quality built for every release

Live demo · sample agents

Know your agent got worse before your users do.

Try the Delta demo
RoleProduct, design and build
TimelineOctober 2026
PlatformWeb app: dashboard, triage, release gate
StatusWorking concept on demo data
Delta eval dashboard with regressions, test results and pass rate across releases
01 the problem

Agents change every week. Nobody can say if they got better.

Teams shipping AI agents change a prompt, swap a model or add a tool, then find out from a customer that the agent now refunds without asking or closes tickets on its own. Unit tests don't catch it, because the same input can pass once and fail the next time.

Most teams test by hand with a spreadsheet of example prompts, or with tools built for engineers that product managers and support leads can't read.

So the brief: run the same test suites on every release, and show the whole team what passed, what broke and whether it is safe to ship.

Delta releases page listing each agent release with its pass rate and gate result
every release gets a score and a ship or block call (demo data)
02 key decisions

Six calls that shaped the product

1

Run every test three times

Agents aren't deterministic, so each test runs three times. 3/3 is a pass, 0/3 a fail, and anything between is flagged flaky instead of hidden in an average.

2

Regressions first, not the score

A pass rate can hold steady while one critical test breaks. The dashboard opens on what passed last release and fails now.

3

Name the failure, not just the fail

Every failed run is tagged: wrong tool, hallucinated fact, overstepped authority, unsafe output, broken format, timeout or lost context. Patterns show up across releases.

4

Authority boundaries are a suite

"Never delete without consent" and "respect the £50 approval limit" are tested like any feature. An agent should know where its authority ends.

5

A gate the team sets

Rules like "block if pass rate drops 5 points" or "block on any critical regression" are editable, so the ship call is agreed before the release, not argued after.

6

Readable by the whole team

Each test shows the input, what good looks like, the agent's steps and the grader's reason in plain English, so a PM or support lead can judge it without reading code.

03 from failure to fix

Every broken test gets an owner

A dashboard that only shows red doesn't fix anything. Delta puts every failing or flaky test from the latest release into a triage queue with a status and an owner.

Open a test and you see its history across releases, the run trace for each attempt and the grader's verdict, with notes for the root cause and the linked fix.

Delta test detail with history across releases and triage
one test, eight releases, one owner
Delta triage queue of failing and flaky tests
the triage queue for the latest release (demo data)
04 what's in it

One workbench for the release

ScreenWhat it answers
DashboardIs this release safe to ship? Regressions, results, pass rate over time and a suite-by-release heatmap
CompareWhat changed between any two releases, test by test
TriageWho is fixing each failing or flaky test, and where it stands
Failure types and costsWhich kinds of mistakes are growing, and what each release costs to run and how fast it answers
SettingsThe release gate rules, plus new test cases and a "run evals" button for a new release
Delta compare releases page
two releases side by side
05 how I'd measure it

Success is a bad release that never shipped

Delta is a concept running on sample agents, so these are targets, not results.

  • Activation: a team runs its first suite against a real release within a day of signing up.
  • Habit: share of releases that go through the gate before shipping, week after week.
  • Guardrail: regressions caught in Delta versus ones customers report first, and flaky tests kept below an agreed share.
06 how it was built

Built with an AI agent, judged by me

I built Delta with an AI coding agent. I made the product calls: three runs per test, the failure types, authority boundaries as their own suite, what the gate blocks on and how a test reads to someone who doesn't code. It's a fast static web app; triage notes, new tests and gate rules save in the browser.

It's a concept: the three agents, their releases and every run are sample data.

07 what's next

From concept to pilot

  • Connect a real agent and run suites from a release pipeline, with results stored in a database.
  • Graders a team can tune, mixing rule checks with an AI judge and human review.
  • A pilot with two or three AI product teams to test the gate on real releases.
08 what I took away

"An agent you can't measure is an agent you can't trust with authority."

Let's talk

Got a product, a hard problem, or just want to say hi? Send it over. I read every message.

G

George Odiana

Open to product roles, side-project collaborations and good conversations about hard product problems.

⚡ 1
let's make something togetherCONTACTdrop a line