Agent Evaluation — Task Success Rate, Eval Harnesses, Benchmarks
Why agent evaluation is its own discipline
An agent does more than generate text — it calls tools, holds state across many steps, reacts to failures, and makes decisions along the way that nobody scripted in advance. That means the classic model question, “is the answer good?”, stops being enough. The real question becomes: did the system get the job done, and did it do so on a path you can trust? Agent evaluation answers that — a discipline that doesn’t test the language model in isolation, but the whole interplay of model, tools, prompt, and error handling.
The eval harness — the test infrastructure behind it
An eval harness is the infrastructure that runs agent evaluations end to end — not a simple unit test, but a system built for non-deterministic outputs. Two pieces make it up:
- Dataset — a collection of test cases (“goldens”) with a task, context, and an expected result or expected property.
- Metric suite — deterministic checks (regex, tool-call counters, schema validation) plus LLM-based scoring for everything that can’t be checked with a hard rule.
The flow: the harness runs every case through the agent, captures the response and the execution trace (which tools ran when, with which arguments), applies the metric suite, and persists the scores. The key difference from software tests: a unit test checks pass/fail on deterministic code. An eval harness answers questions like “was that tool call justified?” or “did the agent ask for clarification at the right point?” — questions no linter or hook can cover.
Task success rate and trajectory accuracy
The obvious metric is task success rate: the share of tasks the agent brought to a correct end result. On its own, it’s easy to be misled by it — τ-bench (Sierra, 2024) found that even GPT-4-based agents succeeded in fewer than half of all tasks, and when the same task was repeated eight times, only about a quarter of runs landed on a consistent result. That’s where the distinction between pass¹ (success in a single run) and pass^k (success across all of k independent repeats) comes from — the latter measures reliability, not just capability.
Trajectory accuracy matters just as much: was the path to the result reasonable and traceable, or did the agent stumble into the right answer via a risky, wasteful, or unsafe detour? An agent that issues a refund without checking authorization and happens to land on the right amount has a correct task success rate — and a real safety problem. That’s why mature harnesses score the full trace, not just the final message.
The key agent benchmarks
| Benchmark | Measures | Domain | |---|---|---| | GAIA | 466 tasks requiring multi-step web browsing, file parsing, and tool use, graded by humans against reference answers | General assistant | | τ-bench | Multi-turn interaction between an agent, a simulated user, and tools/database, including pass^k consistency | Airline and retail customer service | | SWE-bench Verified | Whether an agent turns a real GitHub issue into a diff that makes the tests pass | Coding agent | | HAL (Holistic Agent Leaderboard) | A unified, cost-aware third-party leaderboard spanning nine to eleven individual benchmarks (including SWE-bench Verified, GAIA, τ-bench Airline, AgentHarm), with a Pareto view of accuracy vs. cost | Cross-domain |
HAL, a Princeton project, stands out because it lifts the eval-harness principle to the leaderboard level: one reproducible harness instead of dozens of incompatible setups with different cost assumptions.
LLM-as-a-judge for agents
For agents, a judge that only reads the final answer isn’t enough — it has to understand the trace. In practice, two judge types are used: an outcome judge, which only checks whether the end result is correct, and a trajectory judge, which reads the whole run and flags unnecessary, risky, or wrongly-sequenced tool calls. The known bias classes from plain LLM evaluation — position, verbosity, and self-preference bias (details in Measuring LLM Quality) — still apply, but longer traces add context-length pressure and “quiet drift” when the judge model itself changes over time.
Offline vs. online evaluation
Offline evaluation runs against a fixed dataset, usually as a CI job before deployment — reproducible and cheap to repeat, but blind to what real users actually ask for. Online evaluation scores agents on real or near-real live interactions, often via agent-as-a-judge at runtime. A 2026 study on interactive agents found online judging reached an average criteria coverage of 0.92 versus 0.56 for classic offline judging — online methods capture more of what actually matters in practice, but cost more and only kick in after the interaction has already happened. In practice, the two complement each other: offline evals as a regression gate in CI, online evals as monitoring for drift no fixed dataset could have anticipated.
How it differs from pure model quality
Measuring LLM Quality scores the language model in isolation — factuality, reasoning, style, measured through evals and public benchmarks. Agent evaluation scores the whole system: tool selection, error handling on API timeouts, state management across many turns, cost per completed task, and safety constraints. The two only correlate loosely. An agent built on a top-tier model can still fail on poor tool design, missing retry logic, or too narrow a toolset — and conversely, a cheaper model with a clean harness and tight tool boundaries often delivers a more solid task success rate than a pricier model without agent-specific polish. Model benchmark scores like MMLU or Arena Elo therefore say little about agent performance — which is exactly why GAIA, τ-bench, and HAL exist as their own yardstick.
FAQ
FAQ
- No. Model benchmarks like MMLU measure the isolated language model. Whether the agent turns that into a working tool chain can only be measured by a dedicated agent eval with real tool calls.
- Whether the path to the result was reasonable and traceable — correct tool order, no unnecessary or risky intermediate steps — rather than just checking whether the final answer is correct.
- A fixed dataset can't anticipate every real user request. Studies show online evaluation covers substantially more of the criteria that actually matter, though it costs more and only applies after the interaction.
- For your own use case, start with a small custom task dataset built from real cases. GAIA, τ-bench, or HAL are useful afterward for benchmarking against other agents.
- Real tool calls, sandbox environments, and often multiple repeats per task for pass^k measurements — that pushes cost noticeably above plain text evals, which is why HAL reports cost explicitly.
Is a strong model benchmark score enough to evaluate an agent?
What exactly does trajectory accuracy mean?
Why isn't offline evaluation enough on its own?
Which benchmark should I start with?
What does agent evaluation cost on top of plain LLM eval?
Conclusion
Agent evaluation isn’t a footnote to LLM evaluation — it’s its own discipline: an eval harness runs tasks against the full agent and scores trace and outcome together, task success rate and trajectory accuracy complement each other, benchmarks like GAIA, τ-bench, and HAL provide comparison points, and offline and online judging cover different risks. Testing only the underlying model means not testing the system that actually reaches the user.
Discover more
AI Workflows by Keyword: How We Make Recurring Routines Enforceable
A typed keyword triggers a fixed AI routine — and every single step must be committed before the next one appears. Why that's the actual trick.
GlossaryEnsemble / Multi-Model Orchestration
Ensemble means combining several deliberately varied LLM runs or models whose findings complement each other. Multi-model orchestration drives these runs via orchestrators with sub-agents, so the union of results is larger than any single run.
EncyclopediaFunction Calling / Tool Use
How an LLM uses tools: define a tool as a schema, the model picks the function and arguments, the result returns to the chat — the basis of every agent.