Hugging Face Releases ProxyConfidence: Enhancing Auditing for Black-Box LLM Agents
By Mr.Xu
Published:
Summary:Hugging Face introduces ProxyConfidence, a novel method for auditing black-box LLM agents. By running a low-cost surrogate model in parallel, ProxyConfidence recovers the missing signals in agent decisions and employs complementary evaluation techniques like teacher forcing, request PMI discrimination, and tool-choice competition to detect errors in agent calls. This method requires no access to the agent's internals and significantly improves detection accuracy (AUROC of 0.825) on complex codin
Background and Challenges
When deploying LLM agents, these agents may emit erroneous tool calls, queries, or code, and such issues often surface only after execution. Existing frontier chat APIs hide the model's token probabilities, and the agent's stated confidence barely outperforms random guessing on critical mistakes. Additionally, resampling is ineffective because frontier models are highly repetitive, reproducing the same call across samples.
Core Method of ProxyConfidence
Hugging Face's ProxyConfidence addresses these challenges through the following:
- Parallel Open-Weight Surrogate Model: ProxyConfidence runs a low-cost open-weight surrogate model in parallel, which reads the same context, schema, and proposed action as the agent.
- Complementary Evaluation Techniques:
- Teacher Forcing and Request PMI: Assess the likelihood of each argument value.
- Discriminative Verdict: Judge the call as a whole.
- Tool-Choice Competition: Compare the function against its siblings.
- Trust Principle: Generative likelihood localizes wrong argument values, while the verdict catches holistically wrong calls. When the error type is unknown, an ensemble is the low-regret default choice.
Technical Highlights
- Training-Free: ProxyConfidence requires no training to operate.
- No Access to Agent Internals: The method only requires the agent's output and contextual information.
- Efficiency: The cost of ProxyConfidence is just one prefill pass alongside the tool call.
Experimental Results
On difficult coding tasks, ProxyConfidence achieves an AUROC of 0.825, while the actor's stated confidence is near chance (0.598). The generative readouts outperform the actor's confidence by +0.07 to +0.28 across three different actors. On near-deterministic actors, ProxyConfidence gains +0.14 to +0.19 in self-consistency tests, at 1/K the cost.
Deployment Modes
- Real-Time Review Gating: Escalate the least-trustworthy calls for review, increasing accepted-action accuracy by +0.05 to +0.30 at 50% coverage.
- Confidence Feedback: Return the tool result with the score so the agent adapts its next step, lifting task success on live-execution benchmarks (+0.119 and +0.137, p <= 1e-4) and outperforming a random-value control where step errors are silent (+0.078, p = 0.003).
Industry Impact and Developer Recommendations
ProxyConfidence offers an efficient and practical solution for auditing and error detection in LLM agents, particularly in applications where reliability is critical, such as finance, healthcare, and critical infrastructure. Developers are encouraged to leverage this method to enhance the transparency and reliability of their agents while mitigating the risks associated with errors. It is recommended that developers stay updated on the further development and application cases of this technology to better integrate it into existing agent workflows.
— END —Source: ArXiv AI (cs.AI) (2026-10-07)
Tags: #Hugging Face #LLMs & Foundation Models #Agent Auditing #ProxyConfidence
Community Comments