📈Enterprise Trust in Unreviewed Agent Deploys Drops to 56%
TL;DR
Firms allowing agent production changes on automated evals alone fell from 75% to 56% in a VentureBeat survey. Meanwhile 61% reported at least one customer-facing failure after passing internal tests.
Firms allowing agent production changes on automated evals alone fell from 75% to 56% in a VentureBeat survey. Meanwhile 61% reported at least one customer-facing failure after passing internal tests. The 140-respondent sample is self-selected.
Key Points
Among final decision-makers the figure dropped from 88% to 61%
Firms expecting to keep human approval rose from 20% to 42%
Top evaluation worry: tests do not match real-world outcomes (27%), then lack of explainability (24%)
OpenAI's developer platform appears in 59% of stacks, up from 31% in July
Sample is 140 self-selected respondents, so treat it as directional
Why It Matters
The mood is shifting from 'ship agents unattended' to 'show me the review loop'. Budget is following: human review workflows lead reliability spending at 30%.
Quick Facts
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,566 builders reading daily.