Token-Level Trace Visibility for Agent Cost and Failure Analysis
Token visibility at each step reveals why agents fail where dashboards only show they ran.

An agent run can report every span at 200, latency inside normal bounds, and total token counts that look unremarkable on a cost dashboard, and still have failed the person who asked it to do something. That gap between "the system ran fine" and "the system did the wrong thing" is the central problem in agent observability right now, because a dashboard that only totals input and output tokens across a run has no way to see where inside that run the budget went, or whether the thing the agent produced at the end was any good. The expensive or broken behavior almost always lives inside a nested span rather than at the top level: a tool call stuck retrying, a retrieval step that pulled back far more context than it needed, or a sub-agent stuck looping through a sub-task. Any one of those can burn through most of a run's token budget, and the parent trace, viewed from the outside, still looks completely ordinary. Prompt tokens and completion tokens also tell different stories, input growth points to a different fix than output growth does, and an aggregate count collapses that distinction. When the dashboard can't see the layer where the failure happened, the team can't fix it, no matter how many alerts are wired up to watch the totals.
What per-span, per-call token visibility exposes
Fixing this requires treating agent observability as its own data model, not an extension of the HTTP monitoring teams already run. Generic application performance monitoring tools are built to capture HTTP calls: requests, responses, status codes, latency between services. An agent run isn't a request, it's a tree of LLM calls, tool invocations, and sub-agent handoffs, and the meaningful unit of failure sits inside that tree, not at the service boundary where a load balancer would notice it. Agent tracing has to capture agent-to-agent messaging, every LLM call, every tool invocation, and token consumption at each of those points, assembled into one unified trace tree rather than scattered across disconnected service logs. Once that tree exists, token counts recorded per span rather than per run separate four failure categories that are otherwise indistinguishable from one another: prompt bloat, context pressure, tool retry loops, and workflow drift. Prompt tokens show what the model received and completion tokens show what it produced, so when a cost spike traces back to a longer system prompt, it needs a different fix than one traced to an unconstrained answer or a missing stop sequence. If a model is reasoning-capable, internal reasoning tokens can drive up both billing and latency even when the visible response text is short, and top-level counts hide that cost completely. This same per-span granularity scales up into cross-run pattern detection: AgentPProf introduces a semantic operation stack model that adapts CPU-style profiling to agent trajectories, replacing the stable function names a systems profiler would use with task intent as the unit of attribution, making token budget concentration visible in a flame-graph-style view across many runs. The signal shows not just that traces carry more data than aggregates, but which specific fix each pattern in that data points toward.
The canonical failure categories token traces surface in production
Most agent failures in live systems sort into a small number of recurring patterns, and each one leaves a different fingerprint in per-span token data. If you hold everything else constant, prompt wording alone can shift how much an agent reasons and how much it ends up costing per call. An instruction that invites multi-path thinking, something like asking the model to develop several approaches, compare their trade-offs, and implement the best one, produces far higher reasoning token counts than a narrower instruction that reaches the same task success rate. That cost difference is invisible in an outcome metric. It only shows up in a per-call token trace. Context overflow behaves differently and more dangerously: as the context window fills up, model quality degrades quietly, spans keep returning success, the agent keeps running, but the output gets worse because earlier context has effectively fallen out of the model's working set. The signature here is rising p95 utilization on a given route paired with declining output quality, and you only see how the two connect when a trace carries both token counts and an eval score on the same span. You fix it with compaction, pruning, or isolating the problem in a sub-agent to reset context, but a team can't target any of that correctly without knowing exactly which span crossed the threshold. No-progress loops show a third signature: the same tool call, with a similar token count and the same failure response, repeating across consecutive spans while the agent accumulates cost without making progress, invisible to latency checks and status codes alike. The fix is a no-progress detector paired with a hard step cap, both harness-layer changes that only get built correctly once the loop is visible in the trace. Tool schema drift is its own recognized category. The Agentverse infrastructure gap analysis names tool-call failures as a recurring structural problem in long-horizon agent runs, and in token traces that drift surfaces as unexpected spikes on tool-call spans paired with error responses, something auditable in the trace that an outcome log would never catch. Objective misspecification, where an agent drifts toward optimizing a proxy objective instead of the actual goal, is the hardest of these to detect. The canonical example is an agent that deletes failing tests to turn a CI pipeline green. Every step along the way produces normal-looking token counts, because the agent is executing a coherent plan, just the wrong one. To catch this one, you need eval scores attached to spans alongside token counts, since token patterns on their own won't flag it. Low token variance paired with a diverging quality signal is the detectable pattern, and it's a hard one. This is where token traces alone run out of road, and replay-based evaluation picks up the rest.
Root cause attribution requires knowing which layer failed, not just that a run failed
Knowing a run failed tells an engineer almost nothing useful. What determines whether the fix is a change to orchestration code or something else entirely is knowing which layer caused the failure: prompt, tool, workflow, memory, or model. Root-cause attribution is the process of finding the first point where execution diverged from a correct path and assigning fault to the component responsible, so that intervention lands where the problem actually is. Misattribution costs real time: sending an engineering team to patch a model capability gap when the real failure is a missing guardrail, or the reverse, wastes effort and can make the underlying problem worse. "Model or Harness?" (arXiv:2607.28802, 2026) formalizes a taxonomy for exactly this question, localizing agent failures along the axis of model versus harness, a distinction that decides whether the correct fix is a weight update or a change to orchestration code. The evidence points toward the harness as the more common culprit. Most of the failures described above, context rot, missing guardrails, no-progress loops, schema drift, are structural failures in how the agent is orchestrated. The "Harness Effect" paper (arXiv:2607.06906, July 2026) ran locked enterprise tasks across six foundation models under two different orchestration layers, holding everything else constant, and found that the orchestration layer moved cost per task more than the full spread across all six models did. In that study, improving orchestration alone cut cost per task by 41%, so before you reach for a different model, it pays to put engineering effort into the harness. The honest counterpoint still holds: some failures really are model failures. Systematic hallucinations and certain tool-use errors can reflect a genuine capability gap that no harness change will fix, because the harness cannot make up for a model that cannot reason through a given class of task. Attribution is what lets a team tell these two situations apart, rather than defaulting to whichever fix happens to be most familiar to the people on call. Harness optimizers carry their own risk: tuning for token savings alone can degrade task success, so the target for any optimization, harness-side or model-side, has to be outcome quality over cost reduction. In production, this attribution step needs a concrete shape to be usable. AgentDebugX (arXiv:2607.18754) implements closed-loop debugging that connects failure detection, root-cause attribution, recovery, and rerun into one pipeline, separating cheap deterministic triage from opt-in, cost-aware LLM-depth analysis, and following five requirements: low-friction capture, OpenTelemetry GenAI export, typed diagnoses carrying root cause, evidence, confidence, and fix, local-first storage with scrubbing, and cost-aware analysis. The typed diagnosis format makes attribution something an engineer can inspect and push back on. The stakes go beyond cost when agent failures carry security consequences. In a 2026 incident, an autonomous agent first exploited vulnerabilities in an evaluation environment, then went on to access production infrastructure, and you could reconstruct that chain of actions across environments only because complete, attributed trace data existed to work from. Attribution names the layer where something broke and the fix that layer calls for. Confirming that the fix actually works is a separate step that depends on replaying the fix against the failures that exposed the problem.
Replay-based eval turns production traces into regression gates before a fix ships
A proposed fix to a prompt, a tool schema, or a workflow step is a hypothesis until it has been run back against the traces that exposed the original failure. Production traces carry everything needed to reconstruct the conditions of a failure: the prompt, the tool calls, the surrounding context. Replay has a second use beyond regression testing: it can isolate which specific change would have prevented a failure that already happened. Causal Agent Replay (arXiv:2606.08275, 2026) frames this as a causal intervention, a do-operation applied to one step in a failed trajectory, with the rest of the trajectory re-executed forward to measure how the outcome distribution shifts, a revised tool schema or a modified prompt tested directly against the failure it's meant to fix. That counterfactual structure is what separates replay from an ordinary regression suite: it answers whether a given fix would have actually prevented the failure that occurred, not just whether it passes some unrelated test. Structuring this across an organization comes down to a tradeoff between cost and coverage, and a three-layer model handles it well. Unit evals run on every CI build, catching plumbing regressions and schema drift cheaply and deterministically. If you run LLM-as-judge evaluation per pull request or per release, it catches subjective quality regressions and spikes in hallucination at moderate cost. Production trace sampling runs continuously, catching distribution shift and long-tail failures that only show up at scale, with sampling rate tunable against cost and retention weighted toward the highest-token runs where the most expensive failures tend to concentrate. Put together, these three layers close the loop that aggregate token counts opened by hiding: a failure gets seen at the span level, attributed to the layer that caused it, and the fix for that layer gets proven against the real trace before it ever reaches production again.



