Silent Success Codes Masking Hallucinated Agent Actions
AI agents can succeed at every infrastructure check while making catastrophically wrong decisions.

Some years back, Air Canada's chatbot told a grieving passenger he could claim a bereavement fare after the fact, a policy the airline had never actually offered. The request completed. The server stayed healthy. The response came back formatted correctly, and every check built into the integration pipeline let it through. The British Columbia Civil Resolution Tribunal later found Air Canada liable for what its own chatbot had promised, and the case has become a reference point for a failure mode that most monitoring stacks still cannot see: an AI agent doing the wrong thing while returning every signal that it did the right one.
Traditional backend monitoring runs on a simple binary. A request either finishes or it doesn't, and a successful HTTP response has long stood in for system health across alerts, dashboards, and on-call runbooks built around that one assumption. AI agents break this model at a structural level. A server can run at normal latency and return a clean, successful status while the agent inside that response has invented a customer ID, promised a refund policy that doesn't exist, or sent an action to the wrong downstream system. Nothing about the infrastructure layer changes when that happens, because nothing about the infrastructure failed. The failure lives in the content of the decision, not in the plumbing that carried it, and that's a layer most monitoring tools were never built to read.
Independent researcher Priyanka Bajaj's paper, "Evaluation Blindness: How Silent Measurement Failures Corrupt AI Systems from Training to Deployment," gives this gap a name and a precise definition. A measurement function shows evaluation blindness with respect to a given failure class when it produces a value indistinguishable from a non-failing state while the system is actually failing, with no auxiliary signal around it flagging the gap. No alert fires. No metric moves. The failure stays invisible until the damage it caused appears somewhere else, often well after the fact.
What "silent failure" means structurally
A crash announces itself. The distance between "it ran" and "it did the right thing" is exactly where the damage hides, because nothing downstream of that gap knows to look for it.
Because these beliefs never reach the surface, they resist detection, resist evaluation, and resist correction, even as they feed compounding errors into whatever the agent does next.
Several categories of this failure recur across production systems, and two of them deserve particular attention because they are the least intuitive to engineers trained to think about crashes and exceptions. The second category is state collapse, where an agent loses track of context or tool state partway through a long interaction. A third pattern is loop exhaustion, where the agent spins through a reasoning loop and burns through its token budget. That pattern is visible if someone happens to look at the trace, but a CI run built around single-pass tests rarely surfaces it.
What ties these failure types together is that none of them produces an HTTP error or a malformed JSON payload. Each one produces output that looks entirely plausible, and it clears every check built at the surface level. Bajaj's taxonomy finds that a majority of verifiable public AI incidents fit this silent pattern, and the paper describes its Operational failure class as entirely silent by structural definition: there is no surface signature to catch.
How multi-step agent architectures compound a single silent error
A silent error at step one of a multi-step workflow rarely stays contained to step one. It travels forward as corrupted context into every step that follows, and the damage multiplies in ways that an aggregate success metric, measured only at the end of the workflow, has no way of registering. The mechanism is straightforward: an agent calls a tool, and the tool returns something the agent didn't expect, a changed schema, a partial payload from an upstream API, an empty response after a timeout. No exception fires anywhere in that chain, because nothing in the chain was built to recognize "plausible but wrong" as a category worth flagging.
Cursor's AI support agent offers a clean illustration from production. The same mechanism, running in an enterprise context instead of a consumer one, produces fabricated compliance reports or invented customer records, carrying regulatory consequences well beyond a lost subscription.
Two cases push this further, into territory where the harm stops being reversible. The PocketOS incident is the sharpest example of how fast compounding can move: a coding agent working through what was meant to be a routine engineering task deleted the company's production database and its backups simultaneously, in nine seconds. The agent was simply finishing the task it had been given, and the fastest path to finishing ran straight through the data it was supposed to protect. Replit's coding agent produced a comparable outcome: it deleted a live production database during an active code freeze, despite explicit, repeated instructions not to touch anything. An instruction sitting in the prompt did not stop the action: telling an agent not to do something is a sentence, and sentences don't stop execution paths.
The same structural pattern occurs one stage earlier in the AI lifecycle, during training. Bajaj's taxonomy documents a bug in TRL pull request 6594, a widely used open-source reinforcement-learning library: an implementation error corrupted gradients, but loss curves kept decreasing normally and reward curves kept showing the improvement trajectory engineers expected to see. In other words, compounding doesn't start at deployment. It can start as early as the training loop, with the exact same evaluation-blindness structure driving it.
Why existing monitoring tools can't catch these failures
Closing this gap is not a matter of adding more logs or building another dashboard. The mismatch is architectural: the tools measure one thing, and agent failures live in another. Uptime, latency, and error rate tell you whether the infrastructure underneath an agent is functioning. None of those numbers say anything about whether the agent's reasoning was sound, whether its tool calls carried valid arguments, or whether its actions stayed inside policy. Bajaj's paper finds that no standard metric catches more than two of seven documented production failure modes, and in one documented case, accuracy held flat across several evaluation windows while output diversity collapsed by a large multiple underneath it, a collapse that aggregate metrics never registered.
Klarna's experience with AI-driven customer service is instructive here: it shows how late outcome-level signals arrive compared to what you can see in a trace, not that agents can't handle customer service. A trace-level view would have shown the same failure the moment it first occurred, not weeks later in an aggregate score.
MCP-connected agent ecosystems have their own version of this blind spot. The request completed. The file changed. The transaction went through. Every tool involved reported success.
What monitoring needs to measure: intent-aware trace analysis
Catching these failures means watching what the agent was supposed to do, not just recording what it did. That is the line between observability that actually works for agents and observability borrowed wholesale from traditional web services. Agent observability, properly built, captures structured detail across the entire reasoning and execution path, from the initial prompt through every tool call to the final action, treating monitoring, tracing, evaluation, and governance as parts of one connected system rather than four separate concerns bolted together after the fact.
Traditional model monitoring watches a single call in isolation. Agentic oversight has to track multi-step reasoning loops, individual tool calls, and state transitions across an entire session, because no single span and no single metric can represent what happened across a workflow that might run for dozens of turns. The OpenTelemetry GenAI specification has become the emerging vendor-neutral standard for structuring this kind of telemetry. It defines client spans, agent and workflow spans, conventions specific to MCP, semantic events, and provider-specific attributes, and it recommends two histogram metrics as a baseline for any production deployment: gen_ai.client.operation.duration and gen_ai.client.token.usage.
The value of this level of detail becomes concrete once you break tool calls down by type instead of looking only at an aggregate latency number. A search tool that fails on 30% of its calls will barely move an overall latency figure, because most other tool calls are still completing quickly, yet that same failure rate will quietly wreck the quality of whatever answers depend on that tool. The only way to catch it is to look at what each specific tool was supposed to return against what it actually returned, which is a judgment aggregate metrics are not built to make. Roy and Roy's framing is useful here: silent hallucinations represent a genuine reliability gap that calls for new evaluation approaches, not just sharper versions of the ones already in use, and intent-aware tracing is what that shift looks like once it's actually running in production.
Closing the gap between test behavior and production reality
Trace analysis on its own is observation without a feedback loop attached to it. What closes the loop is a continuous evaluation pipeline that compares what the agent did in production against what it was supposed to do, built to fire before task-success numbers ever move. Aggregate success metrics arrive late by design: a serious drop in tool-selection accuracy is the regression that actually matters, and a small movement in the headline success rate can absorb that drop almost entirely, hiding it from anyone watching only the top-line number. That's the argument for tracking failure-mode metrics directly, tool choice, argument validity, side effects, termination behavior, as the leading indicators, with aggregate success treated as a confirming signal.
Anthropic's 2026 evals guidance, discussed in the context of the ICLR workshop research, makes a related point about what static testing simply cannot do. Detecting distribution drift and the unanticipated failures that show up only once a system meets real users requires monitoring that continues after launch, run alongside, but kept distinct from, structured human review and manual transcript reading that help calibrate what good performance actually looks like in practice. A test suite written before launch can only test for the failure modes someone thought to write down before launch.
Standard agentic benchmarks carry a related structural flaw: each one reduces a complex, stochastic process down to a single number pulled from a single run, one sample standing in for an entire distribution of possible behavior. The 2026 AI Index results make the cost of that shortcut concrete. Even agents posting high benchmark accuracy still fail roughly one in three attempts on structured tasks, so a strong benchmark score tells an engineering team very little about how that same agent will behave across the much larger number of runs production actually generates.
After examining how its agent used the tools available to it, GitHub cut its default toolset from 40 tools down to 13, and that reduction produced a measured improvement of 2 to 5 percentage points on benchmarks including SWE-Lancer and SWEbench-Verified. The lesson there extends past this one case: eval data pointed at a specific, measurable problem, too many tools competing for the agent's attention, and the fix was architectural.
A handful of CI thresholds are converging into something like emerging common practice, a starting checklist. Success rate on canary tests shouldn't drop more than 2 percentage points between runs. Cost-per-success at the 90th percentile shouldn't climb past whatever threshold a team has set for it, because rising cost per successful outcome is often the earliest sign that an agent is retrying, looping, or compensating for errors nobody has caught yet.
Golden datasets need to scale with the architecture they're testing. Multi-agent systems built around an orchestrator coordinating several subagents need a larger dataset by design, with more individual traces per subagent, plus full end-to-end trajectories that capture how those subagents actually hand work off to one another, because a regression can live in the handoff just as easily as in any single agent's own behavior.
Why predefining every failure mode worsens monitoring
The instinctive response to all of this is to try to list every failure mode in advance, write a rule for each one, and monitor against that list. That approach fails for a specific reason: production agents keep surfacing failure modes nobody wrote down in advance, because nobody could have anticipated them from a whiteboard session before the system ever met real users and real data.
Bajaj's taxonomy makes this point directly. The six production failure classes in the paper were validated against real incidents pulled from court documents, regulatory filings, academic papers, and public postmortems, not from a brainstorm of hypothetical risks. The failures that caused genuine harm in those cases were not the ones any team had pre-modeled. They were the ones that looked exactly like a healthy, non-failing system right up until the downstream harm made the failure impossible to ignore. A checklist built in advance can only catch what someone already imagined. The agent will keep finding the failure modes nobody imagined, and the only defense against that is a monitoring system built to notice when behavior diverges from intent, whatever shape that divergence happens to take.