The Agentic Harness

LLM-as-Judge Eval Design for Agentic Task Outcomes

Fixing how LLM judges evaluate multi-step agent trajectories instead of single outputs.

Contributing Editor · · 14 min read
Cover illustration for “LLM-as-Judge Eval Design for Agentic Task Outcomes”
Eval & Validation · September 29, 2026 · 14 min read · 3,156 words

Agentic task evaluation is breaking on a design flaw that most teams haven't named yet: judges built for single-output tasks get pointed at multi-step trajectories, and the assumptions baked into that judge pattern simply don't hold. The fix requires deeper design choices about what the judge sees, at what resolution, and against what kind of reference. LLM-as-Judge Eval Design for Agentic Task Outcomes.

Why single-output judge patterns break on multi-step agent trajectories

The urgency here isn't theoretical. According to LangChain's State of AI Agents report, 57% of organizations now have agents running in production, and quality is the top reason deployments stall LangChain 2026 State of AI Agents report. That's a lot of teams staring at agent output that looks fine on the surface and shipping it anyway, or worse, holding back a system that actually works because the eval can't tell the difference.

An agentic run actually looks like this under the hood: planning, tool calls, retrieval steps, sometimes a handoff to a sub-agent, all chained together across a trajectory that might run dozens of turns deep. The final answer is just the last artifact in that chain.

The mismatch that follows is structural. An agent can call every tool with perfect syntax, correct arguments, sound reasoning at each step, and still fail the task the user actually asked for. Or it reaches the right answer, but by a route so wasteful or so far outside the intended plan that calling it a "pass" hides a real problem. Exact-match grading, the workhorse of single-output eval, cannot see either failure mode.

An agent retrieving data via SQL versus via an API may produce identical outcomes, yet exact-match trajectory grading marks one wrong, because the space of valid execution paths is exponentially larger than for single-step LLM tasks.

There's a second failure mode sitting underneath the first, invisible in the transcript, which makes it arguably worse. Per DAREBench, judges can infer success from outputs that merely sound plausible, without ever confirming the underlying execution actually happened. An LLM judge reading a trace can award a passing grade when a tool call never fired, when the task quietly timed out, or when a required artifact was never produced at all. It's pattern-matching on fluency, the same failure mode language models exhibit everywhere else, just transplanted into the one place teams trusted it least to matter.

The actual design problem centers on what that judge needs to see, at what granularity, and against what reference. Everything that follows in this piece is an attempt to answer that.

The three-layer evaluation stack and the role of LLM judges within it

Agent evaluation resolves into three layers, and confusing them is where most teams lose the thread Anthropic analysis of millions of agent interactions. Outcome metrics ask a binary question: did the agent hit the final goal state or not? That's useful as a triage signal, telling you something broke, but it says nothing about where or why, and stopping there leaves engineers debugging blind.

Trajectory metrics go a layer deeper, scoring each intermediate step: was the tool choice right, were the arguments correct, did the reasoning hold together, did the steps happen in a sensible order. This is where the actual diagnostic signal lives, and this piece spends the most time on this layer because most teams under-invest in it.

System metrics sit alongside both: token usage per task, latency, how often tools get called, whether the agent recovers gracefully from an error. These matter commercially even when the agent is functionally correct, because an agent that solves the task but burns an unsustainable number of tokens doing it cannot ship into production regardless of how clean its trajectory looks.

The three layers function as a diagnostic sequence Anthropic analysis of millions of agent interactions. Check outcome first to confirm something failed. Move to trajectory to locate where. Then attribute at the component level to find which part of the system is responsible.

Not every layer needs an LLM in the loop, either, and getting this wrong wastes both money and trust in the eval. Deterministic checks, things like whether the tool name matches, whether argument types are correct, whether the call sequence follows the expected order, don't need a judge at all. They need exact matching logic, full stop. LLM judges earn their keep at the layer above that: reasoning quality, plan coherence, whether an argument choice made sense given the context, whether a handoff between agents preserved the information it needed to. None of that reduces to a string comparison. Once a team knows which layer a given judge call is serving, the next question follows naturally: what evidence does it need in front of it to do that job well?

What the judge needs to see: decomposing trajectory evidence before scoring

A raw agent trace is a tangle of reasoning text, tool call records, argument payloads, memory reads and writes, and output artifacts, all interleaved in a format no judge prompt was designed to parse holistically. Handing that to a judge and asking "how'd it do?" is asking it to do information architecture and quality assessment at the same time, and it will do both badly.

Production observability has to capture tool selection, tool arguments, model responses, memory reads, memory writes, state transitions, and decision branches as distinct, extractable fields before a judge ever sees the trace. That's the substrate the entire eval depends on.

The TRAJDEBUG framework builds this out explicitly, constructing multi-granularity trajectory views before scoring even starts. The evidence gets structured at different resolutions before the judge prompt gets written, not after. Structure first, prompt second, is the order most teams get backward.

AgentErrorTaxonomy pushes the decomposition further, splitting a trajectory into four operational modules before judging: memory, reflection, planning, and action Anthropic analysis of millions of agent interactions. Attributing each trace segment to its module before the judge ever reads it narrows the scoring surface considerably, since the judge is now answering "was the planning step sound" instead of "was the whole run good," which is a much easier question to answer reliably Anthropic analysis of millions of agent interactions.

Tool-call evidence deserves its own structuring pass. Separate the call record itself (tool name, arguments, return value) from the reasoning that led to it. Surface argument correctness as its own scoring target rather than folding it into a vague "did this step go well" judgment. And flag retries, loops, and recovery attempts explicitly, so a judge can tell the difference between an agent deliberately backtracking to correct course and an agent stuck in a runaway loop that happens to look like deliberation from the outside.

Multi-agent systems add a layer that a lot of eval setups miss entirely. CollabEval distinguishes three separate evidence types: each agent's sub-trajectory measured against its own local task spec, the quality of the information handoffs between agents, and the system-level outcome Anthropic analysis of millions of agent interactions. A judge that only sees the final output has no way to score that middle layer, the handoffs, at all Anthropic analysis of millions of agent interactions. Handoff failures produce a correct-looking final answer built on a broken relay race underneath it.

The practical upshot: before anyone writes a judge prompt, define a structured trace schema that makes each of these evidence types extractable and separately passable. Skipping this step is how teams end up rewriting judge prompts for months without ever fixing the underlying blindness.

Structuring judge prompts for trajectory scoring: rubrics, granularity, and reference anchoring

A single holistic "was this good?" prompt fails for the same reason a single holistic trace fails: it conflates distinct failure modes that need distinct rubric criteria. Reasoning quality, tool selection, argument correctness, and plan adherence are separate dimensions, and grading them together produces a score that can't tell anyone what to fix.

The fix is decomposition before verdict. Have the judge break the trajectory into atomic claims or steps first, then score. This is the same logic behind factoid-level scoring in text evaluation, just carried over into trajectories. Adobe's ground-truth-as-code framework decomposes both the agent's response and the ground truth into atomic claims, then scores precision, recall, and accuracy over the matched claims, a method that produced a 29% improvement in Matthews Correlation Coefficient over a holistic natural-language baseline Skill-based Agentic Evaluation for Real-time Data Science Tasks. That's not a marginal gain. It's the difference between an eval that correlates with human judgment and one that mostly doesn't.

Format matters here too, in a way that's easy to overlook. A judge that scores a prose answer differently than a table for the same underlying facts is measuring presentation, not correctness. Rubrics need to target the factual claim layer underneath the surface form, or the eval ends up rewarding whichever output happens to look more like what the judge expects to see.

Reference design is where the problem resurfaces, and there is rarely one correct trajectory. Valid paths diverge at every branching point in a plan, and anchoring the judge to a single reference trajectory penalizes every valid alternative that doesn't happen to match it.

Three reference strategies handle this differently Anthropic analysis of millions of agent interactions. An outcome-anchored reference defines the end state the task must reach, so the judge asks whether the final artifact satisfies the goal rather than whether the trace matches a golden path. A constraint-based reference flips the frame again, enumerating what must not happen, invalid tool calls, missing required steps, argument type violations, rather than prescribing an exact sequence of what must happen. And for agents working against live data, an executable reference, ground-truth-as-code, encodes the expected answer as a runnable function executed at eval time Skill-based Agentic Evaluation for Real-time Data Science Tasks. That keeps the reference stable against data drift, and the function signature effectively becomes a schema contract that fails loudly the moment an upstream API changes shape Skill-based Agentic Evaluation for Real-time Data Science Tasks. A judge with no real reference is actively misleading.

Failure verdicts need the same evidentiary discipline. TRAJDEBUG's approach requires verbatim evidence for both the specific wrong commitment the agent made and the specific reference condition it violated. Instruct the judge to cite the exact step before it's allowed to issue a failure verdict. That single constraint does more to kill hallucinated judge verdicts than almost any other prompt engineering trick.

Granularity is the last design lever, and it's a tradeoff, not a universal answer. Step-level scoring is the highest resolution and the most attributable, but it's expensive, and it's the right call when debugging a specific failure category. Sub-goal-level scoring matches how task decomposition actually works in practice and lines up with CollabEval's staged approach. Task-level scoring is coarse, but it's exactly right for regression monitoring in CI, where the question is "did this get worse," not "why".

Known judge biases in agentic evals and engineering strategies to counter them

Long trajectories give bias more surface area to work with. The longer and more layered the input, the more room there is for distraction, anchoring, and presentation effects to warp a verdict that should have been about substance.

Explicit framing biases include position effects, where a judge favors content presented earlier or later in a long trace regardless of merit, verbosity bias, where longer traces or longer reasoning text get rated higher independent of actual quality, and authority-style instruction framing, where the phrasing of the prompt itself nudges the verdict. Implicit presentation biases produce these effects: bandwagon effects, sentiment bias where a confident-sounding wrong answer scores higher than an uncertain-sounding right one, model-name visibility where a judge shifts its score once it infers which model produced the trace, and plain distraction from irrelevant steps cluttering the transcript.

Hallucinated correctness, described in the opening section, is, in this light, just a specific instance of verbosity and confidence bias: the judge infers success from reasoning that sounds right, without ever confirming a tool actually executed or an artifact actually got produced.

Mitigations exist, and they're mostly mechanical rather than clever. Randomize step order in the presented trace when order shouldn't affect the verdict. Strip model-identifying metadata before the trace reaches the judge. Require a structured scorecard with a verdict per dimension before any summary score gets issued, which forces decomposition instead of a single holistic gut check. Require verbatim evidence citations for any failure verdict, the same discipline TRAJDEBUG builds in. And wherever the criterion is genuinely binary, tool name correctness, argument type matching, use deterministic checks instead of spending judge calls on questions that don't need semantic reasoning to answer.

Calibration deserves treatment as an ongoing discipline. Measure the gap between judge scores and human expert ratings on a regular cadence, and correct for it, because the agent's task distribution shifts over time and a judge calibrated against last quarter's failure modes drifts quietly out of alignment with this quarter's. Taxonomy of known biases (per arXiv:2604.16790).

Where static LLM judges hit a ceiling: the case for agent-as-judge and neuro-symbolic approaches

Root cause attribution is where pure LLM judges run into a wall that prompt engineering doesn't get you past. Experiments show sole LLM-based approaches produce unreliable, incomplete diagnostic results even at the frontier: strong models top out around 18.15% accuracy on failure attribution datasets arxiv.org DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents. That's a different order of problem than trajectory scoring.

The Who&When dataset, an ICML 2025 Spotlight built from 127 multi-agent systems with fine-grained failure annotations, makes the gap concrete Who&When dataset (ICML 2025 Spotlight) DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents Anthropic analysis of millions of agent interactions. The best methods reach 53.5% accuracy identifying which agent was responsible for a failure, but only 14.2% accuracy pinpointing the specific step where it happened, and some methods perform worse than random guessing at that step level Who&When dataset (ICML 2025 Spotlight) DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents Anthropic analysis of millions of agent interactions. A judge that's right about the "who" barely half the time and right about the "where" only one time in seven isn't a diagnostic tool yet Anthropic analysis of millions of agent interactions.

The reasons trace back to what a static judge fundamentally cannot do. It can't execute code or query a database to independently check whether a claimed result is actually true. It reads a trace as a sequence of text, unable to trace causal dependencies across steps. And it can't reliably distinguish a failure that originated at step 2 from a symptom that merely surfaced at step 8, which means the verdict it produces might be pointing at the smoke instead of the fire.

Agent-as-judge puts a second agent into the same environment the evaluated agent worked in, letting it run code or query the database to verify results directly rather than trusting whatever the first agent reported. ⟦⟧ That agent-judge can give step-by-step feedback, flag which sub-goals actually failed, and build a narrative of performance rather than a single number. It costs more to run, and it's worth that cost specifically for process-oriented tasks where the execution environment is reachable and the diagnostic payoff justifies a second agent's compute.

Neuro-symbolic approaches take a different route, trading holistic judgment for structured representation. AGENTSCOPE abstracts agent behavior into a Reasoning-Action Graph that encapsulates both the reasoning and the action steps, then reasons about correctness through formally defined invariant violation conditions instead of an LLM's impression of the trace DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents. TRAJDEBUG performs evidence-grounded error-lifecycle tracing, running error trigger detection, error state classification, and critical attribution in sequence, evaluated against TrajErrBench, a benchmark of 486 manually annotated failed trajectories drawn from Tau2Bench and SWE-Bench Pro TRAJDEBUG (EMNLP 2026 Findings). Newer causal graph work, treating attribution as a graph traversal problem rather than a text-scoring problem, is showing up under names like "From Flat Logs to Causal Graphs" and AgentTrace DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents.

The design decision that falls out of all this is fairly clean: a static LLM judge is the right tool for trajectory quality scoring at run time, while neuro-symbolic or agent-as-judge methods are the right tool for post-failure root cause attribution, where the actual goal is figuring out which layer caused the break. Using the wrong one for the job is how teams end up with either an expensive judge that can't diagnose anything, or a diagnostic system too slow and costly to run on every production trace.

Connecting judge output to harness layer attribution

Knowing a run failed tells nobody what to fix. A useful judge verdict names the layer responsible, and the research is fairly blunt about the current state of the art here: existing methods mostly succeed at locating the step where a failure surfaced, but fall short of diagnosing the actual root cause, whether that's a badly written prompt, a broken task decomposition, or a logical deadlock the agent talked itself into.

The NexAU framework gives that attribution a concrete target to land on Anthropic analysis of millions of agent interactions. It exposes seven orthogonal component types as explicit, fixed mount points: system prompt, tool description, tool implementation, middleware, skill, sub-agent configuration, and long-term memory Anthropic analysis of millions of agent interactions. That decoupling is what makes component-level observability possible in the first place, because each failure pattern maps cleanly onto a single component class instead of smearing across the whole system Anthropic analysis of millions of agent interactions.

A related question causes this pattern: did the failure originate in the model or in the surrounding harness? An interaction-centric taxonomy from Raj et al. (2026) works through exactly that distinction, and it's the frame judge prompts should be built around. A verdict of "the agent failed" is not an answer. A verdict that identifies prompt ambiguity, tool schema drift, memory retrieval failure, a workflow loop, or handoff loss is.

For a judge prompt to be attribution-ready, it has to produce three things: a verdict at the step or sub-goal level rather than only the task level, so the failure can actually be located in the trace; a named responsible layer rather than a bare quality score; and verbatim evidence from the trace supporting that attribution Anthropic analysis of millions of agent interactions. Anything short of that gives an engineer a grade without giving them a lead.

Even a well-attributed verdict shouldn't ship straight into a fix. A judge that flags a prompt as the failure cause should trigger a change, but that change belongs in front of historical traces before it goes live, because editing a prompt or a tool description without replaying it against real failure cases is shipping on faith. Treating trajectory-level eval as a genuine CI artifact, written early and run continuously rather than bolted on after a production incident, is what turns judge output from a postmortem exercise into an actual engineering discipline.

Sources

  1. Skill-based Agentic Evaluation for Real-time Data Science Tasks
  2. Chapter 8: Agent Evaluation for LLMs: How to Test Tools, Trajectories, and LLM-as-Judge | by Vinod Rane | Medium
  3. DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents
  4. Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering

More in Eval & Validation