myibrahim.cloud

AI / ML · Tuning & Eval

Evaluating LLM apps: past the demo

The gap between 'looks good in the demo' and 'works at p99 in production' is where most LLM apps die. Here's how to measure your way across it.

A founder I know shipped a customer-support copilot that demoed beautifully. The CEO showed it to investors. Investors funded the round. Three weeks after launch, the team noticed the bot had been confidently making up product features that didn't exist. They had no way to detect this until customers started complaining.

They had no eval set. They had a demo.

This is the article that would have saved them.

What an eval set is#

An eval set is a fixed list of (input, expected behavior) examples that you run your LLM application against, on every change, to see if it got better or worse.

It's the LLM equivalent of a unit test suite. The mechanics are different — outputs are open-ended, "expected behavior" is fuzzy — but the role is the same: a regression detector.

The minimum viable eval set#

Start with 30–50 examples. Yes, that's enough.

  • 20 happy-path queries — the things the bot is supposed to handle well. Each with a description of what a "good" answer looks like (don't write the exact output; write criteria).
  • 10 edge cases — ambiguous queries, queries about things outside its scope, queries with adversarial framing.
  • 10 must-refuse cases — queries it should say no to. PII extraction, off-topic, jailbreak attempts.

Build this once. Run it on every prompt change. The discipline of running it weekly catches regressions immediately.

The four kinds of eval#

Different metrics work for different LLM tasks:

1. Exact-match (or regex match). The cheapest. Works for anything with a deterministic right answer — extracting a date, classifying intent, returning structured JSON. If the output should match a schema, validate the schema. If it should contain "Cairo," check for "Cairo."

2. String-similarity to reference. Works for translation, summarization, paraphrase. ROUGE, BLEU, BERTScore — all flawed but useful as relative metrics. ROUGE moves; that's what matters.

3. LLM-as-judge. Use a stronger model to grade the output. Done right, this correlates well with human judgment for most tasks. Done wrong, it's a wasteful coinflip. (See below.)

4. Human review on a sample. Always do this. Pick 10 random outputs per release, have a person rate them. The only way to catch the failure modes the automated metrics missed.

LLM-as-judge: the right way#

The naive version: "Rate this answer 1-10. Output the number."

Don't do that. The numbers cluster bizarrely, the judge tends to score everything 7–8, and slight prompt changes flip the rankings.

The right way:

JUDGE_PROMPT = """
You are evaluating whether the assistant's response correctly answers
the user's question, based ONLY on the provided knowledge base.

User question: {question}
Assistant response: {response}
Knowledge base passages: {passages}

Evaluate using these criteria:
1. Does the response answer the question? (yes / partial / no)
2. Is every claim in the response supported by the knowledge base? (yes / no)
3. Does the response avoid making up product details? (yes / no)

Respond as JSON: {{"answers": "yes|partial|no", "supported": "yes|no", "no_hallucination": "yes|no"}}
"""

Now your judge produces categorical answers, which:

  • Are easy to aggregate (% of responses that scored "yes" on each axis).
  • Are robust to small prompt changes.
  • Map cleanly to human-perceptible quality.

Run the judge with temperature=0. Sample 3–5 times and majority-vote if you want extra reliability. Judge with the most capable model you have — Opus judging Sonnet output beats Sonnet judging itself.

LLM-judge / human-rater agreement (%)
  • Numeric 1-10 (no rubric)42
  • Categorical with rubric78
  • Categorical + 5x majority vote86

Reference-free vs. reference-based#

Reference-based: you have a known-good answer. Compare the model's output to it. Easy when you have ground truth. Often impossible.

Reference-free: you score the output on intrinsic properties — does it follow the format, does it use the right knowledge, does it avoid hallucinations. Harder to set up, but the only option for open-ended tasks.

Most LLM-as-judge work is reference-free. The judge prompt encodes the rubric.

RAG-specific evals#

If you're doing retrieval, evaluate retrieval and generation separately:

  • Retrieval recall@k: did the right chunk make the top-k? Need a hand-labeled set of (query, gold-chunk-id) pairs.
  • Faithfulness: is every claim in the answer supported by the retrieved chunks? LLM-as-judge with the chunks in context.
  • Answer relevance: did the answer address the question? LLM-as-judge.

Tracking all three lets you diagnose where the pipeline broke. Bad retrieval shows in recall@k. Bad prompting shows in faithfulness. Bad routing shows in relevance.

Production sampling: the eval set isn't enough#

Your eval set is a fixed snapshot. Production drifts. Sample 1% of real production traffic for human review. You'll discover:

  • Inputs you didn't anticipate (add to eval set).
  • Outputs that pass automated metrics but feel off.
  • New failure modes from changing user behavior.

The eval set + production sampling is the loop that keeps quality from drifting.

CI integration#

Wire evals into your deployment pipeline. Two patterns work:

# .github/workflows/eval.yml
- name: Run evals
  run: |
    python eval.py --model main --eval-set evals/v1.jsonl > main.json
    python eval.py --model PR-${{ github.event.pull_request.number }} \
      --eval-set evals/v1.jsonl > pr.json
    python compare.py main.json pr.json
    # Fails if accuracy regression > 2%, or hallucination rate up > 1pp

The bar for blocking a PR should be calibrated. Strict enough to catch real regressions, loose enough to allow noise.

What to do when scores stagnate#

After a few rounds of prompt iteration the eval scores stop moving. Three things to try:

  1. Expand the eval set. Maybe you've memorized the test. Add 50 new examples.
  2. Try a different metric. Maybe ROUGE plateaued but human-rated quality is still improving.
  3. Try a different model. Sometimes the prompt is fine and you're capacity-limited.

The point isn't a perfect score. The point is "did this change make things better or worse on a fixed yardstick." The yardstick is the discipline.

Further reading#

  • Hamel Husain, "Your AI product needs evals" — the canonical, opinionated, painfully accurate guide. Read it twice.
  • Eugene Yan, "Evaluating LLM applications" — broad survey with citations.
  • LangSmith, Braintrust, Helicone — eval-oriented tooling. Use one if you don't want to build your own.
  • "Spider-Man" benchmark for hallucinations — small but vivid eval idea worth borrowing.
  • ai
  • evals
  • llms
  • testing
  • production
Need this built? I build ai product or mvp projects for clients worldwide. Tell me about yours.