The Agentic Harness

Regression Suite Construction From Agent Failure Traces

Catch AI agent failures hidden in production traces before they hide in averages.

Senior Writer · · 10 min read
Cover illustration for “Regression Suite Construction From Agent Failure Traces”
Eval & Validation · October 3, 2026 · 10 min read · 2,276 words

A support agent team ships a prompt change, watches the aggregate pass rate tick upward, and calls the release clean. Weeks later it turns out one cohort of refund requests has been failing quietly the entire time, hidden inside an average that looked fine. The fix starts with where regression test cases come from in the first place, since that gap between what the aggregate reports and what a specific slice of users actually experienced is the problem this piece works through. Synthetic test cases come from a developer's mental model of what might break. Production traces come from the system meeting real inputs under real conditions, and they record failures that actually happened rather than ones somebody guessed at.

The distance between those two sources of truth is large, and it's measurable. Research on AgentEval found that end-to-end outcome checks, the kind that only look at whether the final answer passed or failed, achieve 2.17 times lower failure detection recall than step-level evaluation built around the structure of the task. The paper states that prevailing evaluation practices "systematically mask" the intermediate failures that make up most of the real-world error budget. A single pass or fail at the end of a run tells a team almost nothing about which of the steps along the way actually went wrong.

A multi-step agent run touches planning, retrieval, tool selection, function calling, JSON formatting, and final generation, in that order. A tool schema change can leave every HTTP response green while the agent calls the wrong function. These failures are visible only if someone tests the step where they happened, not just the step where the run ended.

The counterargument is that production traces are messy, incomplete, and hard to label at scale, which is true, but it describes a structuring problem rather than grounds for falling back on hand-written test cases.

Why cascading failures make the step that surfaces an error the wrong place to look for its cause

The step where a failure becomes visible in a trace is rarely the step that caused it. Building a regression case against the visible symptom catches that one incident and nothing else, because the same underlying cause can resurface through a different downstream path the next time the system runs.

Research on TRAJDEBUG, from Tsinghua University and Tencent Hunyuan, frames this with precision. The paper's point is sharp: the critical error is not necessarily the first one that occurred, and it is not necessarily the one closest in time to the moment things visibly broke. A team debugging from the end of the trace backward, looking for the step right before the crash, can land on the wrong cause.

AgentEval quantified what that mistake costs. Modeling the dependencies between steps as a directed graph, rather than scoring each step in isolation, added 34 percentage points to root cause accuracy over flat step-by-step evaluation using identical judges and identical rubrics. That's a marginal improvement from a better prompt to the judge model, but it's also evidence that the structure connecting steps carries most of the diagnostic information, and that throwing away the connections between steps throws away the ability to find where a failure actually started.

For regression testing, this has a direct and costly implication. A test case anchored to the symptom step, the point in the trace where the bad output finally appeared, will pass cleanly the next time the same underlying cause expresses itself through a different downstream path. That is not a rare event. The practical fix is to attribute the failure to the layer that actually caused it before extracting anything resembling a test case from the trace. Until that attribution is done, there's no way to know where the regression case's anchor point actually belongs.

Attributing a failure to the layer that caused it before building the test case

Root-cause attribution has to happen before a single assertion gets written. It determines which layer the regression case targets, what the assertion checks for, and whether fixing the problem at that layer will actually close the gap the trace revealed. Skipping this step and writing a test case against whatever looks broken on the surface produces a case that's anchored to the wrong thing, and a wrong anchor is worse than no test at all, because it creates false confidence.

The literature on real-time detection and repair in agent systems describes a four-way split that gives teams their first real decision point when they sit down with a failed trace: is the failure in the model, the harness, the environment, or the grader. Each bucket calls for a completely different kind of intervention. An environment failure, where an external API returned bad data or a dependency was down, is also outside the scope of a harness-level regression case. The failures that belong in the regression suite described here are harness failures: problems in the prompt, the tool schema, the workflow structure, or the agent's memory handling. Those are the layers a team actually controls and can test deterministically.

Attribution is not a clean science, and it should be presented as such. Where a trace leaves that ambiguous, a human reviewer needs to look at it before the case gets promoted into the suite. That review step is not optional overhead; it's what keeps the suite from encoding a misdiagnosis.

What each failure layer looks like in a trace

Each failure category leaves its own signature in a trace, and recognizing that signature is what lets a team route the failure to the right fix instead of patching the wrong layer. This taxonomy isn't theoretical speculation about how agents might fail. These categories describe what actually happens in deployed systems.

Planning errors appear in a trace as wrong tool sequences, loops that never terminate, or steps that silently get skipped. The regression case for a planning error belongs at the planner prompt or at a cycle detector, not scattered across the individual tool calls that followed from a bad plan, since those downstream calls were only doing what a flawed plan told them to do. The research points to three structural weaknesses that account for most agent runs going wrong this way: no exit condition, which lets the agent loop or drift indefinitely; no fallback handling, which causes the agent to halt the moment a single tool call errors out; and no format specification, which produces output that's substantively correct but structurally unparseable by whatever consumes it next.

Tool schema drift is its own dangerous category, precisely because it tends not to announce itself. The model keeps following the contract it learned during training or fine-tuning, while the system around it has quietly moved to an updated contract. A call can return HTTP 200 while the agent has picked the wrong function. Tool failures split further into schema mismatch, which covers malformed or truncated payloads and is the most common and most dangerous class, and semantic garbage, data that's valid and well-formed but wrong in a way no schema validator could ever catch.

Retrieval errors look different in a trace: stale chunks or context that doesn't match the query, surfacing downstream as a drop in groundedness scores or as claims in the final answer that nothing in the retrieved context actually supports. The regression case here belongs at the retriever's filtering logic and at the grounding assertion, not at the model call that came afterward and simply worked with whatever bad context it was handed.

Reasoning errors, meaning hallucinated intermediate steps or inferences the input doesn't support, call for a hallucination score or a carefully calibrated LLM-judge rubric as the primary check. Safety and policy violations carry their own distinct trace signature and deserve a separate cohort with evaluators built specifically for that purpose, rather than being folded into general quality checks.

The scale of the problem across these categories is not small. Each of these failure types maps to a specific place in the trace and a specific evaluator, which is what makes it possible to build a regression case that targets the right thing.

Extracting a regression case from a trace: the fields that make it reusable

A regression case pulled from a trace needs to carry more than the input that triggered the failure and the bad output that resulted. If it carries only those two things, it will catch that exact incident and nothing else, and the entire value of mining production traces was supposed to be catching the class of error, not the single occurrence of it.

A reusable case needs, at minimum, the raw input: the user request or task exactly as it arrived in production, unedited and uncleaned. The case needs the retrieved context or the tool-call arguments that a correct run would have produced, so an evaluator can check grounding and schema compliance against something real rather than against the flawed output the trace actually contains. It needs the failure layer label produced during attribution, prompt, tool, workflow, memory, model, or environment, since that label is what tells the suite which evaluator to attach. It needs a cohort label, tied to product area, customer tier, tool route, or prompt version, so that regressions can be caught at the level of a specific slice of traffic rather than buried in an aggregate. And it needs error propagation metadata, a record of which downstream steps were affected by the failure, drawn from the same dependency structure AgentEval formalizes as a directed graph.

The cohort label is what makes a per-cohort delta the real release gate, instead of an aggregate pass rate that can hide exactly the kind of damage described at the start of this piece. A candidate release can raise the aggregate pass rate and still deserve rejection if a release-critical evaluator or cohort crosses its threshold, which is what happened to the support agent team whose overall pass rate improved while its refund cohort's ToolSelectionAccuracy score dropped underneath it.

Generalization has to be a design goal from the start. A regression case should be built so that it catches the same class of failure arriving through a different input, not simply configured to replay the one incident that produced it. And once a baseline set of cases has been run and recorded, the dataset needs to stay fixed. Every candidate harness runs against the same immutable rows, and the delta between baseline and candidate becomes the decision signal a team actually trusts.

How LLM-generated assertions lock in bugs

The fastest way to build assertions for a regression suite at scale is to hand an LLM the failed trace and ask it to generate the check. It's also the fastest way to permanently encode the bug the suite was built to catch. Research presented at ICST 2025 by Konstantinou and colleagues demonstrated this directly: assertions an LLM writes for a test suite tend to reflect the current implementation, bugs included, rather than the behavior the system was actually supposed to produce. The model looks at what happened and writes an assertion confirming that it happened, which is a different task entirely from writing an assertion that checks whether it should have happened.

The consequence for a regression suite is severe. A suite full of assertions generated this way will pass confidently the next time the same class of error occurs, because the assertion was never checking for correct behavior. It was checking for a repeat of the original mistake. A team relying on that suite gets a green build and a false sense that the system is behaved, right up until a customer notices otherwise.

The fix here is procedural rather than technical. Every assertion extracted from a trace needs human review before it gets promoted into the regression suite, and that review has to check the assertion against what the system was supposed to do, grounded in a spec or a product requirement, not against what the trace happened to show. This is not a step a team can automate its way around.

The Harness Continual Learning framework formalizes a version of this discipline, and it's useful to understand even outside its original context. A Continual Optimizer proposes candidate changes to the harness based on feedback gathered after execution. A Continual Evaluator then decides whether to commit that candidate, but only after checking three things: whether it improves current performance, whether it preserves performance the system already had, and whether it holds up against an anchor set, a collection of previously observed cases with their raw inputs and success criteria, rerun under both the deployed harness and the candidate harness. That separation between the component proposing a change and the component approving it is the same separation a human reviewer provides when checking an LLM-written assertion against the actual spec.

The attribution work done earlier in the process pays off again here. A reviewer who already knows a failure was attributed to the tool layer can check the proposed assertion against the correct tool contract directly, instead of trying to infer correctness from the observed output alone. Attribution turns review from a guessing exercise into a lookup.

One more calibration check belongs in this process. The agent evaluation literature has found that LLM judges tend to favor answers that are long and sound confident, independent of whether those answers are actually correct. A judge that reports high confidence in its own verdict needs to be checked against its real accuracy rate before a team trusts it to gatekeep a release. Skipping that check means trusting a scorer that can be fooled by the same kind of fluent, confident, wrong output the regression suite was built to catch.

Sources

  1. AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking
  2. TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories
  3. Real-Time Detection and Repair of LLM Agent Failures
  4. Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures

More in Eval & Validation