🔬Proxy Confidence: Auditing Black-Box Agents at 0.825 AUROC
TL;DR
A cheap open-weight surrogate runs beside a black-box coding agent and scores its tool calls from log-probabilities. It reaches 0.825 AUROC on hard coding tasks versus 0.598 for the agent's own stated confidence.
A cheap open-weight surrogate runs beside a black-box coding agent and scores its tool calls from log-probabilities. It reaches 0.825 AUROC on hard coding tasks versus 0.598 for the agent's own stated confidence.
Key Points
Submitted October 2, 2026 by Yikai Zhao, Saurabh Pandey and Pradeep Kumar Misra
Beats stated confidence by +0.07 to +0.28 AUROC across three other actors
Self-consistency gains of +0.14 to +0.19 at 1/K the sampling cost
Confidence feedback lifted task success by +0.119 and +0.137 on live-execution benchmarks
Why It Matters
Asking an agent how sure it is does not work well. A small second model watching the logprobs gives you a usable gate for human review without paying for repeated sampling.
Quick Facts
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,558 builders reading daily.