Your agent passes 80% of your tests. That sounds like a solid B. Now imagine a user who asks it to do three things in a session. If each attempt succeeds 80% of the time, the chance that all three work is about half. Most teams discover this from support tickets, not from their test suite, because they measured the wrong number.
This guide shows how to evaluate an agent properly: what to measure, how to grade it, and how to build a small evaluation harness in Python that runs without any API key. Everything here follows Anthropic's own guidance on agent evals, and every number in it comes from code we ran.
Why testing an agent is harder than testing code
A normal function returns the same output for the same input, so one passing test means something. An agent does not. It samples from a model, chooses tools, reacts to what those tools return and may take a different path each time. Running a test once tells you what happened once.
Anthropic's engineering team published a guide on this in January 2026, "Demystifying evals for AI agents". It gives a vocabulary worth adopting, because it makes the problem concrete:
- Task: "a single test with defined inputs and success criteria".
- Trial: one attempt at a task. Because outputs vary between runs, you run several.
- Grader: logic that scores some part of the agent's performance. A task can have several.
- Transcript: the complete record of a trial: outputs, tool calls, reasoning and intermediate results.
- Outcome: the final state of the environment when the trial ends.
- Harness: the code that runs tasks, records every step, grades the results and adds them up.
The number that matters: pass@k versus pass^k
When you run a task several times, there are two very different questions you can ask, and the guide names both:
- pass@k measures "the likelihood that an agent gets at least one correct solution in k attempts". It rises as k grows. It fits situations where someone can pick the best of several tries, such as generating code that a test suite will check.
- pass^k measures "the probability that all k trials succeed". It falls as k grows, because consistency is a harder bar. It fits any agent a customer relies on: they don't get to retry until it works.
Our position is simple: for any agent that acts on behalf of a user, report pass^k. pass@k is useful for research and for tasks with an automatic checker, but it flatters an agent that works only some of the time. In the run below, the same agent scores 1.00 on pass@3 and 0.49 on pass^3. One of those numbers describes your users' experience.
To estimate both from n trials with c successes, we use the standard unbiased estimator from OpenAI's 2021 Codex paper ("Evaluating Large Language Models Trained on Code"): pass@k = 1 − C(n−c, k) / C(n, k). The matching estimate for pass^k is C(c, k) / C(n, k), the chance that k trials drawn from your n all passed.
Grade the outcome, not the path
The most common mistake in agent tests is checking that the agent called specific tools in a specific order. Anthropic's guide warns against exactly this: such tests are "overly brittle", and "it's often better to grade what the agent produced, not the path it took." An agent that reads the calendar before checking the user's preferences, rather than after, is not wrong. An agent that books the wrong day is.
So graders should inspect the final state: the record in the database, the file on disk, the email in the outbox. Three kinds of grader exist, and the guide is clear about their trade-offs:
- Code-based graders (exact checks on the final state) are fast, cheap, objective and reproducible, but brittle to harmless variation. Use them wherever the outcome can be checked by rule.
- Model-based graders (an LLM with a rubric) handle open-ended output and nuance, but are non-deterministic and cost money. Use them for tone, completeness or quality of written answers.
- Human graders are the gold standard, but slow and expensive. Use them to calibrate the model-based graders, not to grade every run.
Build it: a small eval harness in Python
The harness below has no dependencies beyond the Python standard library. It runs every task several times, applies every grader to the final state, keeps the transcript of each trial, and reports pass rate, pass@k and pass^k. Create harness.py:
"""A minimal eval harness for agents: tasks, trials, outcome graders, pass@k and pass^k."""
import json
import zlib
from dataclasses import dataclass, field
from math import comb
from typing import Callable
@dataclass
class Task:
id: str
prompt: str
graders: list[Callable[[dict], tuple[bool, str]]] # each checks the final state
kind: str = "regression" # "regression" (should pass ~100%) or "capability"
@dataclass
class Trial:
task_id: str
passed: bool
failures: list[str] = field(default_factory=list)
transcript: list[str] = field(default_factory=list)
def pass_at_k(n: int, c: int, k: int) -> float:
"""Chance that at least one of k trials passes (unbiased estimate from n trials, c passed)."""
if n - c < k:
return 1.0
return 1 - comb(n - c, k) / comb(n, k)
def pass_hat_k(n: int, c: int, k: int) -> float:
"""Chance that all k trials pass (pass^k), estimated from n trials with c passes."""
return comb(c, k) / comb(n, k)
def run_eval(agent, tasks: list[Task], trials: int) -> list[Trial]:
results = []
for task in tasks:
for i in range(trials):
seed = zlib.crc32(f"{task.id}:{i}".encode()) # same seeds every run
state, transcript = agent(task.prompt, seed=seed)
failures = []
for grader in task.graders:
ok, why = grader(state)
if not ok:
failures.append(why)
results.append(Trial(task.id, not failures, failures, transcript))
return results
def report(tasks: list[Task], results: list[Trial], k: int) -> None:
print(f"{'task':<26}{'kind':<12}{'passed':>8}{'pass@'+str(k):>9}{'pass^'+str(k):>9}")
for task in tasks:
runs = [r for r in results if r.task_id == task.id]
n, c = len(runs), sum(r.passed for r in runs)
print(f"{task.id:<26}{task.kind:<12}{f'{c}/{n}':>8}"
f"{pass_at_k(n, c, k):>9.2f}{pass_hat_k(n, c, k):>9.2f}")
failed = [r for r in results if not r.passed]
if failed:
print("\nFirst failure to read in full:")
print(json.dumps(failed[0].__dict__, indent=2, ensure_ascii=False))
Two details matter. First, each trial gets a fixed seed derived from the task and trial number, so a run is reproducible: the same code gives the same report every time, which is what lets you compare a change against the previous version. Second, the report prints one failed transcript in full. That is deliberate, as you will see below.
Write tasks and graders
Now the tasks. To make this runnable without an API key, the example uses a stand-in agent that manages a task list and makes the kinds of mistakes real agents make: sometimes it drops a due date, sometimes it retries a call and creates a duplicate, and on a multi-step request it sometimes completes the wrong item. In your project, replace demo_agent with a call to your real agent that returns its final state and transcript. Create run_eval.py:
import random
from harness import Task, report, run_eval
# --- Stand-in agent ---------------------------------------------------------
# Replace this function with a call to your real agent. It must return the
# final state it produced and a transcript of what it did. This stand-in
# makes the same kinds of mistakes real agents make, at fixed rates.
def demo_agent(prompt: str, seed: int):
rng = random.Random(seed)
tasks, transcript = [], [f"user: {prompt}"]
def add(title, due=None):
tasks.append({"title": title, "due": due, "done": False})
transcript.append(f"tool add_task(title={title!r}, due={due!r})")
if "invoice" in prompt:
add("Send invoice to Acme", None if rng.random() < 0.10 else "2026-10-10")
if rng.random() < 0.05: # retries after a timeout and duplicates the task
add("Send invoice to Acme", "2026-10-10")
if "three tasks" in prompt:
for title in ("Book venue", "Order catering", "Send invites"):
add(title)
target = 1 if rng.random() > 0.35 else 2 # sometimes marks the wrong one
tasks[target]["done"] = True
transcript.append(f"tool complete_task(task_id={target + 1})")
transcript.append("assistant: Done.")
return {"tasks": tasks}, transcript
# --- Graders: check the final state (the outcome), not the exact steps ------
def has_task(title):
def grade(state):
found = any(t["title"] == title for t in state["tasks"])
return found, f"missing task {title!r}"
return grade
def due_date(title, expected):
def grade(state):
due = next((t["due"] for t in state["tasks"] if t["title"] == title), None)
return due == expected, f"{title!r} due {due!r}, expected {expected!r}"
return grade
def no_duplicates(state):
titles = [t["title"] for t in state["tasks"]]
return len(titles) == len(set(titles)), f"duplicate tasks: {titles}"
def only_done(title):
def grade(state):
done = [t["title"] for t in state["tasks"] if t["done"]]
return done == [title], f"done tasks {done}, expected [{title!r}]"
return grade
TASKS = [
Task("invoice-with-date",
"Add a task to send the invoice to Acme, due 2026-10-10.",
[has_task("Send invoice to Acme"),
due_date("Send invoice to Acme", "2026-10-10"),
no_duplicates]),
Task("three-tasks-complete-2nd",
"Add three tasks for the event, then mark ordering catering as done.",
[has_task("Order catering"), only_done("Order catering"), no_duplicates],
kind="capability"),
]
if __name__ == "__main__":
results = run_eval(demo_agent, TASKS, trials=20)
report(TASKS, results, k=3)
Notice what the graders check: that the task exists, that its due date is right, that nothing is duplicated, that exactly the right item is done. None of them care which tool was called first. The two tasks are also labeled differently, following the guide: regression tasks "should have a nearly 100% pass rate" because they cover things the agent already does, while capability tasks "should start at a low pass rate" because they measure what you are still trying to make it do.
Read the results, then read the transcripts
Running python run_eval.py with 20 trials per task printed:
task kind passed pass@3 pass^3
invoice-with-date regression 16/20 1.00 0.49
three-tasks-complete-2nd capability 14/20 0.98 0.32
First failure to read in full:
{
"task_id": "invoice-with-date",
"passed": false,
"failures": [
"'Send invoice to Acme' due None, expected '2026-10-10'"
],
"transcript": [
"user: Add a task to send the invoice to Acme, due 2026-10-10.",
"tool add_task(title='Send invoice to Acme', due=None)",
"assistant: Done."
]
}
Three things stand out:
- pass@3 hides the problem. Both tasks score 0.98 or higher, which looks ready to ship. pass^3 says a user who repeats the invoice task three times has about a 49% chance all three work, and for the multi-step task about 32%. We checked both by hand: C(16,3)/C(20,3) = 560/1140 ≈ 0.49.
- A "regression" task at 80% is a red flag. Something the agent should always do right fails one time in five. That is the first thing to fix, before any new capability.
- The transcript tells you why. The failure above is not vague: the agent called
add_taskwithdue=None, even though the user gave a date. Now you know whether to fix the tool description, the prompt, or the parameter schema. A score alone could never tell you that.
This is why the guide insists: "You won't know if your graders are working well unless you read the transcripts and grades from many trials." Reading transcripts also catches broken graders. A grader that rejects "96.12" because it expected "96.124991" will report failures that are really bugs in your test.
When the outcome can't be checked by rule
Some outputs, such as a reply email or a summary, have no single correct answer. For those, use a model as the grader, with a strict rubric and a constrained output. Anthropic's testing documentation recommends rubrics that are detailed and specific, scores that are simple ("correct"/"incorrect" or a 1–5 scale), and letting the grader reason before it answers. This is the pattern from that documentation:
def build_grader_prompt(answer, rubric):
return f"""Grade this answer based on the rubric:
<rubric>{rubric}</rubric>
<answer>{answer}</answer>
Output 'correct' or 'incorrect' in <result> tags."""
Before you trust a model grader at scale, grade 20 or 30 outputs yourself and compare. The guide's advice is that LLM judges "should be closely calibrated with human experts". If you and the grader disagree often, fix the rubric before you trust the scores.
A starting plan you can follow this week
- Start small, from real failures. The guide recommends beginning "with 20-50 simple tasks drawn from real failures". Every bug report and every embarrassing demo becomes a task.
- Write unambiguous tasks. If two people could disagree about whether the agent succeeded, the task is the problem. The guide notes that a 0% pass rate across many trials with a strong model usually means "a broken task, not an incapable agent".
- Run every task at least 10–20 times and report pass^k for the number of times a real user would repeat the action.
- Keep regression and capability tasks separate. Regression tasks gate releases; capability tasks track progress.
- Read a sample of transcripts every time, not just the summary table.
- Run the suite on every change to the prompt, the tools or the model. A prompt tweak that fixes one task often breaks another, and only a suite catches it.
Bottom line
An agent is ready when it is consistent, not when it can succeed. Measure consistency with pass^k, grade the outcome rather than the path, and read the transcripts behind every failure. The harness above is about 60 lines of standard Python. Wire your real agent into it, write 20 tasks from your last month of bug reports, and you will know more about your agent's reliability by tomorrow than a hundred manual demos would tell you.
Evals pair naturally with the other two building blocks we covered: giving your agent tools with MCP and giving it memory that survives. Both are exactly the kinds of change your eval suite should run against.
Sources: Anthropic Engineering, "Demystifying evals for AI agents" (January 9, 2026); Anthropic's documentation on developing tests (platform.claude.com); Chen et al., "Evaluating Large Language Models Trained on Code" (2021) for the pass@k estimator. Code tested on 2026-10-05 with Python 3.13; the stand-in agent's failure rates are fixed in the code and are illustrative, not measurements of any real model.