---
title: "The Eval Flywheel"
description: "A Malaysian legal research assistant, an eval harness, and what three loops through the data actually revealed."
date: "2026-05-18"
tags: ["ai", "agents", "evals", "rag"]
---

# The Eval Flywheel

There's a line from Karpathy — written in 2022, mostly ignored at the time — that has aged into something close to a law: **"competitive advantage in AI goes not so much to those with data but those with a data engine."** Whoever spins it fastest wins.

The practitioner community has a name for this now. Hamel Husain calls it the *eval flywheel*. Eugene Yan calls it *Eval-Driven Development*. The names vary but the shape is the same. Six steps running in a loop, each one feeding the next.

I've been building a [Malaysian legal research assistant](https://github.com/aishahsofea/ai-legal-tool). The agent has four fixed stages: classify the query, retrieve relevant statute chunks from a database of Malaysian legislation, synthesize a grounded response, then validate every citation before anything goes out the door. I've spun the flywheel several times now, and here's what actually happened.

## The Loop

Every serious eval practitioner converges on the same workflow, with cosmetic variations. Log your pipeline's full traces. Look at them without an agenda. Do error analysis by annotating failures freely, clustering them into patterns, and counting which ones recur most. Write evals that capture the top failure class. Fix the system. Measure. Watch for regressions. Repeat.

What makes this a *flywheel* and not just a checklist is that each turn generates the input to the next. You are back at error analysis after every loop, but now you know more.

Click each node to see what it meant in practice.

> **Interactive Eval Flywheel Diagram:** [view it in the web article](/blog/the-eval-flywheel)

## Three-level Pyramid

The most useful operational framework is Hamel's **three levels of evaluation**. Cost grows by roughly an order of magnitude per level, which dictates how often each runs. The intuition is simple. There's no point paying for an LLM's opinion on a response that already failed a basic check. So cheap deterministic assertions run first, and if any fail, the case stops there. The LLM judge only gets called on responses that cleared every fast gate.

> **Interactive Eval Levels Chart:** [view it in the web article](/blog/the-eval-flywheel)

## Three Loops

Theory is one thing. The build log is another. What follows is what actually happened, with each loop triggered by a score drop and each fix validated by re-running the suite.

<div class="not-prose my-8 overflow-hidden rounded-lg border border-zinc-100 dark:border-zinc-800">
  <div class="flex flex-wrap items-baseline gap-x-3 gap-y-1 border-b border-zinc-100 dark:border-zinc-800 bg-zinc-50 dark:bg-zinc-900/50 px-5 py-3">
    <span class="font-mono text-[10px] tracking-widest uppercase text-zinc-400 dark:text-zinc-600 shrink-0">Loop 01</span>
    <span class="text-sm font-medium text-zinc-800 dark:text-zinc-200">When the Right Fix Is a Scope Decision</span>
    <span class="ml-auto shrink-0 rounded-full border border-red-200 dark:border-red-900/50 bg-red-50 dark:bg-red-950/30 px-2.5 py-0.5 font-mono text-[10px] text-red-500 dark:text-red-400">language_register 0%</span>
  </div>
  <div class="px-5 py-5">
    <div class="mb-4 overflow-hidden rounded-lg bg-zinc-100 dark:bg-zinc-800/60">
      <div class="flex items-center gap-1.5 border-b border-zinc-200 dark:border-zinc-700/50 px-4 py-2.5">
        <div class="h-2 w-2 rounded-full bg-red-400/80"></div>
        <div class="h-2 w-2 rounded-full bg-amber-400/80"></div>
        <div class="h-2 w-2 rounded-full bg-zinc-300 dark:bg-zinc-600"></div>
        <span class="ml-2 font-mono text-[10px] text-zinc-500 dark:text-zinc-400">run_evals.py --smoke</span>
      </div>
      <div class="space-y-0.5 p-4 font-mono text-xs">
        <div><span class="font-medium text-red-500">FAIL</span> <span class="text-zinc-600 dark:text-zinc-300">language_register: 0/5 = 0.0%</span></div>
        <div class="pl-5 text-zinc-400 dark:text-zinc-600">↳ BM query received English-only response</div>
        <div class="pt-0.5"><span class="font-medium text-emerald-600 dark:text-emerald-500">PASS</span> <span class="text-zinc-600 dark:text-zinc-300">citation_existence: 8/8 = 100.0%</span></div>
        <div><span class="font-medium text-emerald-600 dark:text-emerald-500">PASS</span> <span class="text-zinc-600 dark:text-zinc-300">uuid_leakage: 10/10 = 100.0%</span></div>
        <div class="pt-0.5 text-zinc-400 dark:text-zinc-600">Judge: 71.4% — below 80% gate</div>
      </div>
    </div>
    <div class="mb-3 text-sm leading-relaxed text-zinc-600 dark:text-zinc-400">The Bahasa Malaysia cases were scoring 0% on the language assertion because the pipeline was responding in English to BM queries. Clear failure on paper. But before writing a fix, error analysis pointed somewhere unexpected: the entire statute corpus is in English. BM retrieval degrades by design until a BM corpus is ingested. Given what the pipeline had to work with, the model wasn't doing anything wrong.</div>
    <div class="mb-3 text-sm leading-relaxed text-zinc-600 dark:text-zinc-400">This changed the question from "how do we fix the language handling?" to "should we be testing this at all right now?" The answer was no. The BM test cases were removed from the smoke suite and BM support was explicitly deferred to v2, when there will actually be BM content to retrieve against.</div>
    <div class="mb-4 text-sm leading-relaxed text-zinc-600 dark:text-zinc-400">The assertion is still in the code. The capability just isn't claimed yet.</div>
    <div class="border-l-2 border-accent pl-4 text-sm italic text-zinc-500 dark:text-zinc-400">The harness found it either way, but what to do with the finding was still ours to decide.</div>
  </div>
</div>

<div class="not-prose my-8 overflow-hidden rounded-lg border border-zinc-100 dark:border-zinc-800">
  <div class="flex flex-wrap items-baseline gap-x-3 gap-y-1 border-b border-zinc-100 dark:border-zinc-800 bg-zinc-50 dark:bg-zinc-900/50 px-5 py-3">
    <span class="font-mono text-[10px] tracking-widest uppercase text-zinc-400 dark:text-zinc-600 shrink-0">Loop 02</span>
    <span class="text-sm font-medium text-zinc-800 dark:text-zinc-200">The Supervisor Was Calibrated to Claude</span>
    <span class="ml-auto shrink-0 rounded-full border border-red-200 dark:border-red-900/50 bg-red-50 dark:bg-red-950/30 px-2.5 py-0.5 font-mono text-[10px] text-red-500 dark:text-red-400">judge 30% → 80%</span>
  </div>
  <div class="px-5 py-5">
    <div class="mb-4 overflow-hidden rounded-lg bg-zinc-100 dark:bg-zinc-800/60">
      <div class="flex items-center gap-1.5 border-b border-zinc-200 dark:border-zinc-700/50 px-4 py-2.5">
        <div class="h-2 w-2 rounded-full bg-red-400/80"></div>
        <div class="h-2 w-2 rounded-full bg-amber-400/80"></div>
        <div class="h-2 w-2 rounded-full bg-zinc-300 dark:bg-zinc-600"></div>
        <span class="ml-2 font-mono text-[10px] text-zinc-500 dark:text-zinc-400">run_evals.py --mode full # GPT-4.1 trial</span>
      </div>
      <div class="space-y-0.5 p-4 font-mono text-xs">
        <div><span class="font-medium text-red-500">FAIL</span> <span class="text-zinc-600 dark:text-zinc-300">judge: 3/10 = 30.0%</span></div>
        <div class="pl-5 text-zinc-400 dark:text-zinc-600">↳ all failures via FINAL_FAILURE_RESPONSE</div>
        <div class="pl-5 text-zinc-400 dark:text-zinc-600">↳ supervisor Rule 2 firing on every GPT response</div>
        <div class="pt-1 text-zinc-500 dark:text-zinc-500">debug_case.py <span class="text-zinc-400 dark:text-zinc-600"># node-by-node tracer</span></div>
        <div class="pl-5 text-zinc-400 dark:text-zinc-600">router → statute_lookup ✓</div>
        <div class="pl-5 text-zinc-400 dark:text-zinc-600">retriever → 8 chunks ✓</div>
        <div class="pl-5 text-zinc-400 dark:text-zinc-600">synthesizer → response ✓</div>
        <div class="pl-5"><span class="text-red-500">supervisor → Rule 2 FAIL</span> <span class="text-zinc-400 dark:text-zinc-600"># citation regex mismatch</span></div>
      </div>
    </div>
    <div class="mb-3 text-sm leading-relaxed text-zinc-600 dark:text-zinc-400">During a model trial, swapping Sonnet for GPT-4.1 dropped the judge pass rate to 30%. Every single case ended in the pipeline's fallback error response, not because GPT-4.1 got the law wrong, but because a validation rule was rejecting its output before the judge ever saw it. A single-case debug tracer that prints the output of each stage in isolation made this visible in minutes.</div>
    <div class="mb-3 text-sm leading-relaxed text-zinc-600 dark:text-zinc-400">The culprit was a citation pattern check written to match Claude's phrasing: <em>"Section 90A of the Evidence Act 1950."</em> GPT-4.1 writes it differently — <em>"Section 90A(1) states that..."</em> with the Act name appearing earlier in the sentence. The check didn't account for that, so every GPT-4.1 response failed the validation, triggered a retry, and eventually hit the fallback.</div>
    <div class="text-sm leading-relaxed text-zinc-600 dark:text-zinc-400">Extending the pattern check to also accept subsection notation moved GPT-4.1 from 30% to exactly 80%, with all deterministic assertions passing and the judge scoring 8 out of 10. The larger point is that the validation was silently calibrated to one model's citation style. Any future model trial would have hit the same wall. The harness found what a code review wouldn't have.</div>
  </div>
</div>

<div class="not-prose my-8 overflow-hidden rounded-lg border border-zinc-100 dark:border-zinc-800">
  <div class="flex flex-wrap items-baseline gap-x-3 gap-y-1 border-b border-zinc-100 dark:border-zinc-800 bg-zinc-50 dark:bg-zinc-900/50 px-5 py-3">
    <span class="font-mono text-[10px] tracking-widest uppercase text-zinc-400 dark:text-zinc-600 shrink-0">Loop 03</span>
    <span class="text-sm font-medium text-zinc-800 dark:text-zinc-200">The Cost vs. Fidelity Tension</span>
    <span class="ml-auto shrink-0 rounded-full border border-zinc-200 dark:border-zinc-700 bg-zinc-100 dark:bg-zinc-800 px-2.5 py-0.5 font-mono text-[10px] text-zinc-500 dark:text-zinc-400">decision: eval = prod</span>
  </div>
  <div class="px-5 py-5">
    <div class="mb-3 text-sm leading-relaxed text-zinc-600 dark:text-zinc-400">Eval cost was creating friction. With 30+ model calls per smoke run, iteration was expensive enough that you'd think twice before running it. The natural instinct was to swap in a cheaper model for evals.</div>
    <div class="mb-3 text-sm leading-relaxed text-zinc-600 dark:text-zinc-400">The flywheel stopped this. Running evals on a cheaper model while deploying a different one in production isn't an eval. It's a measurement of a different system. Previous testing had already shown the cheaper model would consistently omit citations from its structured output, returning empty lists even when the prose mentioned the right statute. The citation existence check would pass, but the citations would be missing. The eval would be blind to exactly the failure class it was supposed to catch.</div>
    <div class="mb-4 rounded-r-lg border border-zinc-200 dark:border-zinc-700 border-l-[3px] border-l-accent px-4 py-3">
      <div class="font-mono text-[10px] tracking-wider uppercase text-accent mb-2">The compound-probability problem</div>
      <div class="text-sm italic text-zinc-500 dark:text-zinc-400">"A 90% accurate process repeated 5 times is 59% accurate."</div>
      <div class="mt-2 text-sm text-zinc-500 dark:text-zinc-400">Using a slightly different model in evals vs. production compounds this gap in ways that only surface as mysterious production failures with no eval signal to explain them.</div>
    </div>
    <div class="mb-4 text-sm leading-relaxed text-zinc-600 dark:text-zinc-400">The decision was to keep the production model in the eval and address cost by making that model cheaper instead. GPT-4.1 is cheaper than Sonnet. A small routing layer now maps model names to the right provider, so future model trials are a one-line environment variable change, not a code change.</div>
    <div class="border-l-2 border-accent pl-4 text-sm italic text-zinc-500 dark:text-zinc-400">Cost pressure is real and worth solving, but the answer is always to make the right system cheaper, not to measure a different one. The moment you cut the eval model, you've cut the signal.</div>
  </div>
</div>

## Criteria Drift

Shreya Shankar's UIST 2024 paper names the phenomenon precisely:

> To grade outputs, people need to externalize and define their evaluation criteria; however, the process of grading outputs helps them to define that very criteria.
>
> — Shreya Shankar · UIST 2024

I felt this when writing the judge prompt. My initial pass criteria were vague, things like "accurate" and "appropriately hedged." But the moment I tried to write calibration examples, the vagueness collapsed. What does "hallucinated citation" mean exactly? What counts as AI-refusal boilerplate versus a legitimate disclaimer? What's the boundary between refusing to give legal advice and refusing to answer a legitimate research question?

Each of those questions had to be answered explicitly because the judge demanded it. The answers are the product's quality contract, written in concrete terms that a model can apply consistently.

Phillip Carter, after building Honeycomb's LLM judge, put it plainly: *"Seeing how the LLM breaks down its reasoning made me realize I wasn't being consistent about how I judged certain edge cases."* The PRD doesn't precede the eval. **The eval is how you write the PRD.**

## What I'd Do Differently

**Start with 15 hand-graded cases, not 40.** The smoke set should have been the starting point, hand-graded before anything automated was trusted. Starting larger meant trusting automation before the rubric was calibrated.

**Build a single-case debug tracer on day one.** A small script that runs one query through each pipeline stage and prints what came out of each one was only built after the GPT-4.1 failure made the full eval suite useless for diagnosis. It should have been the first tool. Running 15 cases to isolate a bug in one stage is slow and expensive; running one case with full intermediate output takes seconds.

**Version the judge prompt like it's the product spec — because it is.** It deserves the same review rigor as anything else that ships, and a changelog when it changes.

<div class="not-prose my-6 rounded-r-lg border border-zinc-200 dark:border-zinc-700 border-l-[3px] border-l-accent px-4 py-3">
  <div class="font-mono text-[10px] tracking-wider uppercase text-accent mb-2">What compounding actually feels like</div>
  <div class="text-sm text-zinc-500 dark:text-zinc-400">Each loop in this build started with a score drop that <em>already existed</em>. The fix was attempted with measurement, not against vibes. Every pinned smoke case is a promise that this specific failure will never silently reappear. The loop closes. The next spin is cheaper than the last. That's the flywheel.</div>
</div>

The flywheel isn't glamorous. It's test cases, rubrics, error taxonomies, regex fixes, and cost calculations. But it's the part of agent engineering where the work actually compounds.
