MCP Tool Evaluation and Testing in Agent Evals
Dynamic tool discovery breaks traditional pass-fail testing designed for fixed tool menus.

An MCP-connected agent does not call from a fixed menu of tools. It discovers that menu at runtime, from one or more MCP servers, and the menu can differ from one request to the next. That single fact breaks the premise underneath the deterministic test because the test assumes a fixed tool set, the one that says: given this input, the system should call this tool. There is no "this tool" to anchor the test against once the tool set itself is a variable.
Why pass/fail tool-call checks fall short for MCP-connected agents
A functional test tells you a tool works when you call it by name. It does not tell you whether the agent, faced with a catalog it has never seen before, picks that tool correctly out of everything else on offer. Those are different questions, and the gap between them is where production failures live. A server can return valid JSON on every single call, pass every unit test written against it, and still fail the moment it ships, because the agent chose the wrong tool from the catalog it was given. Nothing in a pass/fail harness catches that difference, because pass/fail harnesses are built to check whether a named function executes, not whether an agent reasoning over a list of tool descriptions picks the right one.
This gap has a name. MCP AI agent readiness describes exactly this question: whether a model, reading nothing but tool definitions, picks the right tool and calls it correctly. Readiness is a property of the match between what the tool's description says and what the model does with that description under real, varied phrasing. A tool can be perfectly implemented and still be unready, because the agent that has to select it has no access to the implementation.
The agent's entire picture of an MCP server comes from the output of a single call: tools/list. If the description is vague, the model is vague. Tool description quality is the eval surface itself, not a documentation nicety sitting off to the side of the engineering work, because it is the only information the agent has to act on.
That surface, in practice, is in rough shape. A study examining hundreds of tools across more than a hundred MCP servers found that 97.1% of tool descriptions carried at least one defect, with more than half never clearly stating what the tool actually does. Against that backdrop, a pass/fail check that only confirms a tool executes correctly when called by name is answering a question nobody in production is actually asking. The question in production is whether the agent would have called it at all, and under what phrasing, and in place of what else.
What trajectory-level evaluation measures across an MCP run
The fix is to change the unit being measured. Instead of scoring a final answer, or scoring one isolated tool call, trajectory-level evaluation scores the full run: which tools the agent called, in what order, with what arguments, and whether it read the results correctly enough to move the task forward. A trajectory is the complete record of an agent's decisions across a task.
Toloka's MCP evaluation framework turns that idea into five concrete questions asked of every run. Did the agent choose the right tool at the right time? Did it construct valid arguments for that tool? Did it correctly interpret the data the tool returned? Did it stay inside its policy boundaries? Did its reasoning hold together across multiple steps rather than drifting or contradicting itself along the way? Answering those five questions requires looking at the entire sequence of a run, not a single input-output pair, because a wrong choice at step two can produce a technically correct-looking call at step four.
Salesforce AI Research built an open-source framework, MCPEval (arXiv:2507.12806), around the same premise: that evaluation needs to happen at the level of the full task, automatically, rather than through hand-built benchmarks scored one interaction at a time. MCPEval automates end-to-end task generation and evaluation of LLM agents across different domains, standardizes the metrics used to score them, and integrates directly with the agent's own native tools rather than a separate test harness bolted on afterward. Its results across five real-world domains show something that static, single-call benchmarks consistently miss: performance gaps that are specific to a domain, visible only once you look at how an agent handles a sequence of related tool calls inside that domain's particular structure.
Trajectory evaluation also is not a one-time gate run before a release ships. Toloka runs it in weekly sprints during training or fine-tuning, which lets a team track whether a model is actually improving over time and catch a regression while it is still small, instead of discovering it after it has shipped to users. Treating trajectory evaluation as a recurring measurement, rather than a single checkpoint, is what makes it useful as a management tool and not just a diagnostic one.
The four trajectory dimensions that MCP evals must cover
A complete trajectory evaluation breaks into four dimensions, each of which can fail on its own, independent of the other three: tool selection accuracy, argument correctness, chain efficiency, and result integration. Each one needs a different kind of check.
Tool selection accuracy asks whether the agent picked the right tool out of a catalog it discovered at runtime, using nothing but description text. This is a retrieval problem, not an execution problem, and it fails in two distinct directions. Both trace back to poor tool-description quality in practice, and both appear clearly on a simple confusion matrix, which is how the FutureAGI functional eval frames the split. Adjacent-tool ambiguity is the trap to watch for: when two tools sit close together in embedding space, a small shift in how a user phrases a request can surface the wrong one, and the agent then either invents arguments to fit a schema that doesn't match the task or routes a write operation through the wrong tool. The LLMFunctionCalling template, eval ID 98, covers this in a single call.
Argument correctness asks whether the agent filled in that tool's schema correctly once it was chosen. A schema missing a required array leaves the model with no way to tell which parameters are mandatory and which are optional, so it guesses, producing either an invalid-params error or, worse, a call that succeeds with the wrong arguments and never surfaces as a failure. This layer can be scored without an LLM judge at all, using deterministic metrics: function_name_match, parameter_validation, function_call_accuracy, and function_call_exact_match. The production target is high schema compliance across the full volume of production calls; a drop usually means an upstream server's schema has drifted out from under the agent.
Chain efficiency asks whether the agent is doing the task in a reasonable number of steps or churning through redundant, circular calls that run up latency, token cost, and the surface area for failure, without getting any closer to finishing the task. A chain efficiency ratio, the minimum number of calls the task actually requires divided by the number the agent actually made, gives a working threshold: above 0.7 is workable, below it the agent is likely looping or repeating work it has already done. Loops are detectable mechanically, by linting the trace for the same tool called with the same arguments more than once in a single run, no judgment call required. A large tool surface creates selection noise on its own, and that noise compounds into inefficient chains. Latency budgets make the same point in numbers: in a four-call chain with an eight-second budget visible to the user, the combined tool calls need to stay under roughly 1.5 seconds, and a p95 latency that climbs above the seven-day baseline is an early warning that chain efficiency is degrading.
Result integration asks whether the agent actually used what a tool handed back. A call can return 200 with perfectly valid data and the trajectory still fails, because the agent never references that data in its next turn, drops a key field from it, or replans from the original prompt as though the call had never happened. Three checks catch this: does the agent reference the returned data on the next turn, does the chain actually progress to the next step, and does the final answer cite the correct value from what came back. Failures here get misread constantly as reasoning failures in the model, when they are usually harness failures: the context window management or the result-parsing layer dropped the data before the model ever had a chance to see it. Toloka's approach to this layer combines automated reward scoring with expert human review of flagged and sampled full trajectories, which catches the kinds of failure that automated metrics alone tend to miss.
Three failure categories that repeatedly break MCP trajectories in production
In production, trajectory failures cluster into three recognizable categories: tool schema drift, prompt ambiguity with adjacent-tool confusion, and workflow loops. Each one leaves a specific signal in the trace beyond a drop in final-answer quality.
Tool schema drift happens when a remote MCP server changes its endpoints or parameter types and the agent is still working from a cached description of the old version. The agent reads a description that no longer matches what the tool actually does or accepts, and nothing announces this out loud. A drifted description that undersells a tool's new capability produces a recall problem, where queries that should hit the tool no longer do. The fix is not another round of eval runs against the same stale description. It requires coupling tool registration to automated schema synchronization or to runtime contract verification, so the description an agent reads is never more than one source of truth behind the tool it describes.
Prompt ambiguity and adjacent-tool confusion happen when tool descriptions are vague or overlap with each other, so a small change in how a user phrases a request is enough to pull the retriever toward the wrong tool. The agent simply cannot tell two tools apart using description text alone, because the text does not give it anything to tell them apart with. Context-token bloat makes this worse before the user's request is even read: seven popular MCP servers alone consume 67,300 tokens of context, and a large tool catalog crowds out the model's working context before the actual query gets processed, dragging down selection accuracy across the board. The most subtle version of this failure is out-of-scope misuse: when a tool's description never states what the server cannot do, an agent asked for a capability outside its actual scope will often force the request through the closest available tool instead of declining the request.
Workflow loops appear in the trace as an agent repeating the same tool call with the same arguments inside a single trace, a pattern detectable by linting the trace mechanically, with no LLM judge needed to flag it. AgentDebugX, submitted in July 2026 by researchers at the University of Illinois Urbana-Champaign, the University of Toronto, Google, and Stanford University, structures debugging as a closed loop: Detect, Attribute, Recover, Rerun. The Attribute step is built specifically to find the first point where execution diverged from correct behavior, which separates genuine root-cause attribution from simply reporting the symptom where the failure happened to surface. A February 2026 survey by Wang et al. cites CHIEF as achieving the highest attribution accuracy among reasoning-based methods for this kind of fault tracing.
Layering MCP evals across development, pre-merge, and production
None of this belongs at a single checkpoint in the development cycle. MCP evaluation works as three layers, each with a different scope, a different cadence, and a different cost, and all three need to read from the same underlying trace tree, or the numbers each layer reports will stop agreeing with each other.
At development time, the right checks are structural and deterministic. Validating a tool's schema against its actual runtime signature, checking required fields, type correctness, and enum alignment, runs in under 10ms and catches mechanical errors before an LLM is ever called into the loop. Tools like MCP Inspector, run with npx @modelcontextprotocol/inspector, let a developer interact directly with a server's tool list and schemas during development, confirming that what tools/list actually returns matches what the tool is meant to do, well before that description is ever handed to an agent to reason over.
Before code merges, the checks shift from structure to behavior. This is where the golden-corpus tests for tool selection accuracy belong, along with the deterministic argument-correctness metrics: function_name_match, parameter_validation, function_call_accuracy, function_call_exact_match. These are cheap enough to run on every pull request and specific enough to catch a regression in tool selection or argument handling before it ever reaches a user.
In production, the evaluation has to run continuously against live traffic, sampled rather than exhaustive, since full trajectory review on every single request is not economical at any real scale. Treating these three layers as one connected system, rather than three separate projects, is what keeps a passing pre-merge suite from becoming a false signal once an agent meets the tool catalog it actually has to work with in production.


