Companies that have already seen an AI agent pass internal testing only to fail in front of customers are moving faster than others to remove humans from deployment decisions, according to new research from VentureBeat Intelligence. The July survey of 108 enterprises found that 85% of those burned by such an incident are pursuing a model where agents can push code or change systems without a person’s approval in certain low-risk cases, either now or within the coming year. That compares with 61% of companies that reported no such customer-visible failure. Only 11% of the burned group rejected end-to-end deployment automation for the years ahead, versus 24% of the unburned group.
The finding is counterintuitive because trust in automated evaluation is rising even as evidence of its limits persists. In July, 13% of respondents said they fully trust automated evaluation, up from 5% in June. Meanwhile, the share citing poor alignment between tests and real-world results as their biggest concern fell from 29% to 19% month over month. Yet 49% of respondents said an AI agent or LLM-powered feature that had cleared company testing later created a problem visible to customers, nearly unchanged from 50% in June, and 24% said this had happened more than once.
The gap between confidence and outcomes is stark when comparing affected and unaffected companies. Among enterprises that experienced a test-passing failure, only 4% placed complete faith in automated checks, while 24% of those with no such incident expressed full confidence, a sixfold difference. The data suggests that firsthand proof of a release gate failing does little to slow the push toward autonomy, and may instead reflect deployment maturity, where higher-volume operators both encounter more failures and have the infrastructure to automate releases.
The shift is reshaping the market for evaluation tools, according to Raindrop.ai, an automated agent error monitoring and mitigation platform. Its CTO, Ben Hylak, said Fortune 100 companies are reducing eval sets and deprioritizing maintenance, finding it impossible to fully enumerate failure cases as systems grow more complex with MCPs and subagents. Instead, he said, they are leaning on anomaly and issue detection solutions before and after production.
The July survey drew responses from 108 people at companies with at least 100 employees, down from 157 in June. Of those, 69% described themselves as final AI-buying authorities or influencers, and 63% worked at firms with 100 to 2,499 employees. The industry mix shifted, with technology and software participation falling nine points to 14%, while retail and consumer rose four points to 19%. The findings are directional, as the survey is self-selected, and the burned-versus-unburned splits rely on subgroups of 41 to 53 respondents.
Across the two months, the proportion of organizations reporting at least one test-approved system that disappointed customers held steady at 50% in June and 49% in July, based on 265 total responses. That does not mean half of all agent runs fail, but it does indicate that a passing internal evaluation cannot be treated as proof of production reliability. The report also flagged four month-to-month changes: more respondents expressing complete confidence, fewer naming poor real-world alignment, more choosing integration ease as the decisive buying factor, and Braintrust gaining primary-platform share.




