Evaluating agents

Evaluating an agent means measuring its behavior on a fixed set of tasks under a fixed procedure, so that results can be compared across checkpoints, models or harnesses. Training reward says whether optimization is succeeding against the reward. Evaluation asks what the policy can now do, and for agents the conditions it holds fixed include a harness, a runtime and budgets, each of which can move a score as much as a better checkpoint.

The evaluation protocol

Two scores are comparable only if they were produced the same way. The protocol of an agent evaluation includes:

  • the tasks and the split they come from
  • the harness, its version and its tool set
  • the runtime image and its network policy
  • sampling settings, which are part of the behavior policy
  • budgets: token limits, turn limits and timeouts
  • the number of rollouts per task
  • the scorer, including the judge model if one is used
  • for conversational tasks, the user simulator's model and prompt

A higher token limit lets more tasks finish, and a different harness changes what the model sees, so a report states whether it used a standard harness, which compares checkpoints, or the product harness, which measures what users get. A simulated user is a second sampled policy, so its randomness adds variance of its own, and changing its model can move the score with no change to the agent. Success belongs next to tokens, turns, wall time and cost, because a policy can raise its success rate by spending more of each.

pass@k and its unbiased estimator

pass@k is the probability that at least one of kk independent attempts at a task succeeds, averaged over tasks. pass@1 measures reliability, and larger kk measures coverage, meaning whether the policy can solve the task at all.

The standard estimator, from the HumanEval paper, draws n≥kn \ge k samples per task, counts the cc that succeed, and computes the probability that a random subset of kk contains at least one success:

pass@k=1−(n−ck)(nk)\text{pass@}k = 1 - \frac{\binom{n-c}{k}}{\binom{n}{k}}

The fraction is the probability that all kk chosen samples are failures. The estimator is unbiased, while the plug-in estimate 1−(1−c/n)k1 - (1 - c/n)^k is biased low. With n=10n = 10 and c=3c = 3:

  • pass@1 =1−7/10=0.30= 1 - 7/10 = 0.30
  • pass@5 =1−(75)/(105)=1−21/252≈0.917= 1 - \binom{7}{5}/\binom{10}{5} = 1 - 21/252 \approx 0.917
  • the plug-in estimate for k=5k = 5 is 1−0.75≈0.8321 - 0.7^5 \approx 0.832

Binomial coefficients overflow for large nn, so implementations use a product form:

import numpy as np

def pass_at_k(n: int, c: int, k: int) -> float:
    if n - c < k:
        return 1.0
    return 1.0 - np.prod(1.0 - k / np.arange(n - c + 1, n + 1))

The complementary pass^k, introduced with τ-bench, is the probability that all kk attempts succeed, estimated as (ck)/(nk)\binom{c}{k}/\binom{n}{k}. With c=3c = 3 of n=10n = 10, pass^2 =3/45≈0.067= 3/45 \approx 0.067. It measures the consistency that matters for an agent users run many times. pass@k also differs from best-of-k: pass@k assumes an oracle that recognizes the success, while best-of-k needs a selector, such as a reward model or the agent's own tests, to pick it.

The two ends of the kk axis can move differently under RL. Yue et al. found RLVR models beating their base models at small kk while the base models reached higher pass@k at large kk, a result that longer training and stricter success criteria complicate (RL scaling). Reporting several values of kk shows which end moved.

Held-out tasks and environments

A held-out split of the training distribution tests whether gains transfer to instances that never provided gradient, and a held-out environment, with different tools, state or verifier, tests a stronger form of transfer. Training reward that rises while held-out performance stays flat points to memorized artifacts or reward hacking.

A split is only as independent as its construction. Tasks from the same repository, templates filled with different values, or near-duplicate problems leak across a random split, so splitting by repository, source or template measures transfer more faithfully. The suite, its primary metrics and its sampling budget should be fixed before the run, so the headline comparison is not selected afterward.

Contamination

A benchmark whose tasks appeared in pretraining data measures recall. Wu et al. gave Qwen2.5-Math-7B the first 60% of each MATH-500 problem and found it completed 54.6% of them word for word, against 0% on the newer LiveMathBench. Prefix-completion tests, benchmarks released after the model's data cutoff, and private task sets are the available checks.

What agent benchmarks fix

Public agent benchmarks differ in which parts of the protocol they fix:

BenchmarkMeasuresFixesLeaves open
SWE-bench Verified500 Python repository issueshuman-screened specs and testspublic repositories, so contamination
SWE-Bench Pro1,865 multi-file tasks from 41 repositoriesheld-out and commercial repositoriesharness and budget choice
Terminal-Bench 2.089 command-line tasks, each with its own environmenthuman-written solutions and testsat 89 tasks, a standard error near 5 points at 50% success (binomial estimate)
τ²-benchsupport tasks where agent and user both act on shared statetask generator and constrained user simulatoruser-simulator variance
BrowseComp1,266 hard-to-find facts on the webshort answers checked against a referencethe live web changes and answers can leak
METR time horizonhuman task length at 50% agent successa unit comparable across modelsdepends on the task suite and human baselines

Variance and confidence intervals

An eval score is an estimate, and its uncertainty is often larger than the difference being reported. Treating the tasks as a sample from a larger population, as Miller recommends, gives a standard error for the mean over MM tasks. For M=200M = 200 tasks with one binary attempt each and a success rate of 0.400.40:

SE=0.40×0.60200≈0.035\text{SE} = \sqrt{\frac{0.40 \times 0.60}{200}} \approx 0.035

A 95% confidence interval is about ±1.96×0.035≈±0.068\pm 1.96 \times 0.035 \approx \pm 0.068, or [0.33,0.47][0.33, 0.47].

More rollouts per task help less than expected. With rr rollouts per task, the variance of the mean splits into a part from differences between tasks and a part from sampling within each task:

Var⁡(sˉ)=1M(Var⁡(pi)+E[pi(1−pi)]r)\operatorname{Var}(\bar s) = \frac{1}{M}\left(\operatorname{Var}(p_i) + \frac{\mathbb{E}[p_i(1 - p_i)]}{r}\right)

where pip_i is the policy's true success rate on task ii. Suppose Var⁡(pi)=0.15\operatorname{Var}(p_i) = 0.15 and E[pi(1−pi)]=0.09\mathbb{E}[p_i(1-p_i)] = 0.09, which add up to the total 0.240.24 above. Going from r=1r = 1 to r=4r = 4 shrinks the standard error from 0.0350.035 to 0.0290.029, and r=16r = 16 only reaches 0.0280.028. Most of the uncertainty comes from which tasks are in the set, and only more tasks reduce it. Rollouts per task are still needed for pass@k and per-task difficulty.

Two checkpoints evaluated on the same tasks call for a paired comparison: compute the per-task difference did_i and its standard error Var⁡(di)/M\sqrt{\operatorname{Var}(d_i)/M}. Task difficulty affects both checkpoints alike and cancels. If a new checkpoint goes from 80 to 90 successes on 200 tasks by fixing ten tasks and breaking none, the improvement is 0.05 with a paired standard error of about 0.015, a 95% interval of roughly [0.02,0.08][0.02, 0.08], even though each score's own interval is wider than 0.05.

Reading beneath the mean

An aggregate hides which tasks changed. Per-environment results belong in the report next to any macro average, and difficulty slices show whether a policy improved only where the base model already succeeded occasionally. The fractions of tasks where every rollout succeeds or every rollout fails show whether the frontier of partially solved tasks is moving, which connects evaluation to difficulty filtering. Error and truncation rates belong in the report, because a crash or a slow sandbox is not a wrong answer.

A small, stable set of traces reviewed across checkpoints separates failures that look identical in the score: a reward shortcut, a harness that hid needed information, an environment timeout, a verifier that rejected valid work, or an improvement bought with far more compute.