VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Enterprise Agent Evals Are Failing — and the Response Is Less Oversight, Not More

Half of enterprises ship agents that pass evals then fail customers. Burned organizations respond by removing humans and shifting to live behavioral…

Vector Wire — AI-assisted editorial illustration

Nearly half of enterprises shipped an AI agent that passed its evaluations and then failed a customer — and the organizations that experienced that failure are accelerating toward autonomous deployment faster than those that haven't. ANALYSIS The emerging enterprise response to broken evaluation is not to add human checkpoints but to replace static benchmarks with behavioral monitoring of live systems, a shift visible across survey data, practitioner discussion, and at least one production architecture.

Why it matters

The gap between evaluation confidence and production reliability is not closing. Across 108 enterprises surveyed in July 2026, the share that fully trust automated evaluation nearly tripled, from 5% in June to 13%1. Yet the failure rate held steady: 49% of organizations deployed an agent that passed internal evaluations and then caused a customer-facing failure. Trust is rising while the underlying problem remains unchanged — a divergence that makes the question of *how* enterprises evaluate agents existentially important.

The big picture

The confidence surge belongs almost entirely to enterprises that have not yet been burned. Among those that have experienced a false-confidence failure, only 4% fully trust automated evaluation; among those that haven't, 24% do. But the burned cohort is not retreating to manual review. Among enterprises that have shipped an evaluation-passing agent that then failed a customer, 85% are on the autonomy trajectory — removing humans from the loop or engineering pipelines to permit zero-human deployment within a year. Among those that haven't experienced such a failure, 61% are on that trajectory.

ANALYSIS The data describes a counterintuitive pattern: failure destroys trust in automated evaluation but accelerates the move away from human oversight. That only makes sense if the burned enterprises are concluding that the problem is the *kind* of evaluation, not the absence of human gatekeepers.

A separate survey of 101 enterprises on context infrastructure reinforces the diagnosis. Sixty-eight percent have traced a confident but wrong agent answer to missing or inconsistent business context — not model error — in the past six months2. Among enterprises running a governed semantic layer, the rate of *recurring* context failures is more than twice that of enterprises without one. Governed data layers are not preventing failures; they are surfacing failures that were previously invisible, suggesting that static evaluation misses context-dependent errors by design.

Between the lines

Practitioners are converging on a replacement paradigm: observe the agent's behavior in the live environment rather than scoring it against predetermined answers. Brex CEO Pedro Franceschi described the shift in concrete terms at VB Transform 2026. When Brex deployed the open-source OpenClaw agent into internal roles, its security team rejected the proposal because traditional security models failed3. The response was CrabTrap, an open-source HTTP proxy that monitors all outbound network traffic between the agent's container and the internet. "People talk a lot about agents, but I think 'agents' is a terrible name. It's this Silicon Valley concept that doesn't really mean much," Franceschi said. Brex's framing — treating agents as "virtual employees" with email addresses and Slack access — demands security and evaluation at the network layer, not the code layer.

The same logic appears in practitioner discussion around SQL-writing agents. The failure mode that concerns builders is not a query that errors out but one that "executes fine and returns real rows — just the wrong ones. Wrong join, wrong filter, stale understanding of the schema"4. Static eval sets with prewritten golden answers break down because the correct answer changes as the data changes. Major platforms — LangSmith, Braintrust, Arize — do not ship live data verification out of the box; online scoring generally falls back to reference-free LLM-as-judge approaches.

ANALYSIS Across these strands, the pattern is consistent: static, pre-deployment evaluation cannot keep pace with agents operating in dynamic environments. The enterprises furthest along are shifting evaluation from a gate before deployment to continuous monitoring during deployment — watching what the agent *does* rather than what it *scores*.

The vendor landscape reflects the transition. The share of enterprises running no dedicated evaluation tooling fell from 17% in June to 12% in July. Fifty-six percent intend to adopt a new or replacement evaluation platform within twelve months. Human review workflows remain the top planned investment at 31%, but production observability is close behind at 30%. Ease of integration rose from 27% to 39%, becoming the top selection criterion for evaluation vendors. Enterprises are not abandoning evaluation — they are shopping for a different kind, one that integrates into production infrastructure rather than sitting upstream of it.

What's next

The 67% of organizations that already allow or are actively engineering zero-human-in-the-loop deployment for low-risk agents represent the near-term frontier. With 49% still experiencing post-eval failures and 68% tracing wrong answers to context failures rather than model error, the pressure to build runtime behavioral monitoring — network-level proxies, live-data verification, production observability — will intensify. The evaluation market's next twelve months will be shaped by whether platforms ship these capabilities natively or whether enterprises, like Brex with CrabTrap, continue building them in-house.

CORRECTIONS: none for this article · this piece updates automatically as the story develops · corrections policy & trail →