The Agentic Harness

Retraining vs Harness Iteration for Agent Performance Improvement

Finding the real cause of agent failure matters more than retraining the model.

Staff Writer · · 11 min read
Cover illustration for “Retraining vs Harness Iteration for Agent Performance Improvement”
Harness Engineering · October 8, 2026 · 11 min read · 2,532 words

Retraining is rarely the lever that fixes a deployed agent. Most performance failures trace back to the harness, the prompt structure, tool definitions, workflow logic, and memory handling that surround the model, and fixing those is faster, cheaper, and far more precisely targeted than touching model weights. This piece lays out why attribution comes first, what the evidence says about harness iteration versus retraining, and how to run the fix as a repeatable process.

The difficulty of fixing agent failures without layer attribution

If an agent fails at step forty, the failure rarely started there. A planning error introduced early in a trajectory distorts everything downstream: the agent's memory carries the bad assumption forward, its reflection step reasons from a corrupted premise, and each subsequent tool call compounds the distortion until the task collapses. Agents stitch together planning, memory, reflection, and tool-use modules into a single pipeline, so one root-cause error doesn't stay contained. It propagates through every decision that follows, and the point where the run visibly breaks is almost never the point where the trouble started.

Research on this cascade dynamic makes the propagation pattern precise. TRAJDEBUG, published by researchers at Tsinghua and Tencent Hunyuan, found that failed trajectories typically contain multiple local errors, and those errors don't all matter equally. Some get silently repaired by a later step. Some sit there harmless, never touching the outcome. Only a subset actually contributes to the final failure, and that subset is neither the first error an engineer spots in the log nor the one sitting closest to the moment everything fell apart. An engineer scanning a trace for the obvious break point will often land on a symptom, not a cause.

That distinction carries real consequences for what happens next. The same research draws a sharp line between what different failure types should trigger: a model failure should inform post-training objectives, a harness failure should guide a redesign of the harness itself, and an environment or grader issue should trigger a repair to the benchmark or evaluation setup. Picking the wrong lever wastes engineering time and leaves the actual problem untouched. Fixing agent performance starts with knowing, concretely, which layer produced the failure, and that question is harder to answer than it first appears.

A structured failure taxonomy mapping error categories to fix surfaces

Sorting agent failures by the module responsible isn't a bookkeeping exercise. It tells an engineer whether to open the harness code or retrain the model, and for the overwhelming majority of production failures, the answer points to the harness.

TRAJDEBUG's framework requires grounding each error in one of several sources: a conflict with the task instructions, a break from the trajectory history, a misread of environment feedback, or a flaw in the agent's own reasoning. The first three of these are harness phenomena by definition. They describe what context got built, what the system told the model, and what state got passed from one step to the next, none of which touches the model's weights.

Three structural harness failure modes account for most agent execution failures, so check these before you suspect the model itself. Tool schema drift is one: the model keeps calling a tool the way it learned to call it, while the system has since started enforcing a different contract, with new required fields or a changed argument shape. That's a configuration mismatch, not a reasoning failure, and no amount of retraining fixes a tool schema that changed after the model was trained. The second mode covers prompt and workflow design choices: missing exit conditions that let an agent run past the point where it should stop, absent fallback handling for when a tool call fails, and output formats that were never pinned down precisely enough for the model to follow consistently. The third mode is workflow loops, where a failed tool call, a truncated context window, or runaway iteration traps the agent in a repeating cycle that standard application monitoring tools can't see, because they weren't built to track agent-specific state across steps.

Scale AI's research on root-cause attribution formalizes this same logic at the system level: fault gets assigned to the component actually responsible, whether that's the model, the harness, the environment, or the grader, and that assignment is what directs the fix to the right part of the system. Systematic attribution, pinpointing whether a failure originates in the prompt, tool definition, workflow logic, or memory layer, is the step that makes a reliable fix possible. Platforms like Moda automate this diagnosis by analyzing production traces to surface which layer broke, turning what used to be manual detective work into a tractable engineering workflow. Once a failure lands in one of these harness-layer categories, tool schema drift, a missing exit condition, a runaway loop, the path forward becomes concrete: adjust the prompt, redefine the workflow, or tighten the tool contract. Trace-driven tools can generate and propose those exact changes, so an engineer can act on the taxonomy directly, not just understand it.

The difficulty of locating the root cause inside a production trace

Diagram: Root-Cause Recovery Rates on Real Failed Trajectories. Visualizes: Show the accuracy contrast between three methods for identifying the root-cause step in failed agent trajectories, all measured on the LongRCA benchmark (1,140…

Even if you know the categories, finding the actual failing step in a real trace isn't easy. In production trajectories that run long, often spanning dozens of tool calls and intermediate reasoning steps, the evidence needed to judge any single step is scattered across instructions given far earlier, observations retrieved pages back, and context that has already partially fallen out of the window by the time an engineer examines the failure. Add to that the fact that a single failed trajectory usually contains several local errors at once, each with a different downstream effect, and the step an engineer's eye lands on first is frequently the wrong one.

The accuracy numbers on this problem are sobering. On the LongRCA benchmark, a set of 1,140 human-annotated failed trajectories spanning five domains, the strongest of five baseline methods, a system called ECHO running on DeepSeek-V4-Flash, recovers the exact root-step only 13.2% of the time. RCTA, a purpose-built method, raises that figure to 24.1%, nearly double the baseline, but three out of every four reference root causes still go unrecovered. These are not casual estimates; they reflect what happens when automated systems, built specifically for this task, are pointed at real failed trajectories.

Scale AI's work on this problem, published under the name Continual Search, explains part of why the numbers land where they do. One-shot LLM judges have to diagnose a failure from a trace in a single pass, so they often settle on a plausible answer early and stop looking, and critical evidence can sit unexamined deeper in longer traces. So reliable attribution needs an iterative search that keeps challenging the standing answer turn after turn, rather than accepting the first explanation that fits the visible symptoms.

The practical consequence is straightforward: human review of execution logs at production scale isn't feasible, and a wrong attribution wastes the engineering cycle spent chasing it. Automated, trace-aware attribution isn't a convenience here. Even with a sound taxonomy in hand, finding the real root step in a long, multi-error trace by manual inspection is unreliable, because human reviewers settle on plausible diagnoses early and miss evidence buried deeper in the log, the same failure mode the research describes in automated judges. Tools built for agent troubleshooting can iterate over these traces systematically, so they catch root causes that manual review passes over, and they check proposed fixes against historical data before anything ships to production.

Harness iteration versus prompt tweaking

Harness iteration means something more specific: you don't just edit a system prompt until the outputs look better. It's the systematic, evidence-driven editing of every layer of the scaffolding wrapping a model, and the breadth of that scope is exactly where its leverage comes from.

The harness is the external system that organizes how an agent calls tools, manages its context and internal state, controls how it interacts with its environment, and mediates each step of a multi-step trajectory. It sits entirely outside the model weights, so you can edit it in full without retraining anything. Four components make up the direct targets for this kind of iteration. Prompts cover system instructions, task framing, output format specifications, exit conditions, and fallback handling. Tool schemas cover argument definitions, how a contract gets versioned over time, and how errors and retries get handled. Workflow structure covers the sequencing of steps, how sub-agents get orchestrated, and the conditions that tell a loop when to stop. Memory and context management cover what an agent keeps, what it compresses, and what it discards as it moves through a task.

Research from LIFE-HARNESS tested harness iteration directly against prompt-only optimization, and the prompt-only approach delivered only modest gains by comparison. The gap reflects something true about agentic tasks generally: performance depends not just on the words in the initial prompt but on how the entire runtime mediates tool calls, actions, feedback, and the many steps that make up a trajectory. A better prompt can only move a system so far when the loop termination logic is broken or the memory layer is dropping context it needs.

What separates harness iteration from ad hoc patching is discipline in how changes get made. That validation step is what turns an edit into something an engineer can trust.

The empirical case that harness iteration outperforms model-level changes for deployed agents

Diagram: Harness Iteration vs. Prompt-Only: A 120% Performance Gap. Visualizes: Show a magnitude contrast between two approaches to agent performance improvement: prompt-only optimization versus full harness iteration.

The research record backs a clear ordering: systematic harness iteration beats prompt-only optimization, and in most deployment situations, it beats model retraining too. The two approaches aren't rivals that cancel each other out; they work together, with harness iteration doing most of the heavy lifting.

LIFE-HARNESS quantifies the gap precisely. Measured against prompt-only optimization, full harness iteration produces an average relative improvement of 120% in Pass@1. A gain of that size says something about where the real performance headroom sits in agentic systems: not in the wording of the initial instructions, but in how the runtime handles tools, actions, and feedback across the full length of a trajectory.

A natural objection follows: if the model later gets better through its own training, does harness work become redundant? Harness-R1 answers this directly. Direct supervised fine-tuning on the agent raises performance on the unmodified target, and a harness engineer built specifically for that stronger, post-SFT agent raises performance further still, by several additional percentage points on top of the SFT gain. So the harness engineer keeps producing value as the underlying model improves, and it doesn't go obsolete once the model gets stronger. Harness-first means the harness is the lever to pull first, with model-level work extending those gains.

A real limitation deserves stating. Harness optimization changes the scaffolding around a model, not its weights, so the gains stay tied to that specific harness at deployment time. A team running a single general-purpose agent across many use cases faces a genuine choice between one shared harness and a set of specialized ones, and each carries its own routing and orchestration overhead. That tradeoff is the actual constraint teams need to plan around, not a reason to skip harness work and default to retraining instead.

When retraining is the right answer

Retraining earns its place only when a failure is demonstrably rooted in the model's own capability or its ability to generalize, not in how the harness is configured around it. The conditions that qualify are narrower than most teams assume going in.

Scale AI's root-cause attribution research sets the gate clearly: model failures, once separated out from harness, environment, and grader failures, are the category that should actually inform post-training objectives. Retraining without that separation is effort aimed in the wrong direction, because a model can look like it's failing on reasoning when the real problem is a tool schema it was never told had changed.

Post-training carries its own specific risk when it ignores the harness. If the tool environment shifts even slightly from what they trained on, models post-trained against a harness built with low design effort can drop sharply in performance. Retraining locks a model to the exact harness state it saw during training, which makes it fragile against the schema changes and workflow adjustments that happen routinely in production.

Three conditions mark the cases where retraining is the right call: root-cause attribution consistently points to the model's own reasoning or capability across many separate traces, not to tool schemas, prompt structure, or workflow design; the task domain has shifted so much that no harness edit can close the gap between what the model knows and what the environment now demands; or the harness has already been iterated to the point of diminishing returns, and attribution keeps landing on the model layer regardless. Even once one of these conditions holds, you still run the sequence harness-first, then retrain on whatever evidence remains after that cleanup. Retraining on noisy, unannotated failure traces without first clearing out the harness-level problems produces a model that inherits those same structural flaws, baked into its weights.

Running harness iteration as a repeatable engineering workflow

Harness iteration works as a disciplined loop: collect traces, attribute the failure, make a scoped edit, validate the fix, and keep monitoring. Treating it as ad hoc patching defeats the purpose.

Start by collecting and triaging production traces. Sample live traffic continuously instead of leaning on synthetic benchmarks, because they miss the actual distribution of failures a deployed agent runs into. The target is a representative set of failed trajectories pulled from real usage, not a curated set of the easy cases that happen to be convenient to analyze.

Attribute each failure to its responsible layer before touching anything. Use the taxonomy already established, tool-schema failures, prompt-structure failures, workflow-loop failures, memory failures, and separate all of those from genuine model-layer failures. Continual Search's findings apply directly here: iterative, multi-pass attribution over a long trace substantially beats a one-shot diagnosis, so build that iteration into the triage process rather than accepting the first plausible answer a review turns up.

Make the edit narrow once attribution points somewhere specific. If the trace evidence implicates a tool schema, fix the tool schema. Resist editing the entire prompt because it's the easiest thing to open; scope the change to the exact component the evidence names.

Validate the edit against historical traces before it goes live. Replay the modified harness against a golden set of recorded traces, including runs that previously passed, so you can check that the fix resolves the attributed failure without breaking something that worked before. TRAJDEBUG's application studies back this step up directly: feeding actionable diagnostic output from trace analysis back into the system measurably improves downstream agent success.

Keep monitoring production after the fix ships, for recurrence of the same failure and for new failure modes the change might expose. Observability here needs to run continuously, not as a periodic check, because trace analysis surfaces new failure patterns in production before engineers formalize them into a dashboard metric. That closes the loop from a production signal, through a scoped harness edit, to a validated improvement, and running that loop on a steady cadence is what keeps a deployed agent from decaying quietly while no one is converting its failures into test coverage.

Sources

  1. TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories
  2. Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
  3. LongRCA Bench: Root-Cause Localization in Long-Horizon Agent Trajectories
  4. Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories

More in Harness Engineering