ZICQ
中 Log in / Sign up
Newsroom LLMs & Foundation Models #JudgeStealer #Model Extraction #Large Language Models #AI Safety #Evaluation Protocols

ArXiv Proposes JudgeStealer: Efficient Extraction of LLM Judging Capabilities

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:The ArXiv team introduces JudgeStealer, a novel model extraction framework designed to efficiently replicate the judging capabilities of large language models (LLMs) across various evaluation protocols. By leveraging cross-protocol agreement, dynamically selecting semantically diverse inputs, and applying score smoothing and multi-protocol review techniques, JudgeStealer significantly improves query efficiency and extraction accuracy. Experimental results demonstrate that JudgeStealer outperform


Key Breakthroughs

The ArXiv team introduces JudgeStealer, an innovative model extraction framework designed to address the critical challenges of replicating the judging capabilities of large language models (LLMs). The key technical highlights of this framework include:

  • Cross-Protocol Agreement Utilization: JudgeStealer leverages the strong agreement between different evaluation protocols to acquire pointwise scores and transform them into pairwise and listwise supervisions, eliminating the need for additional victim queries.
  • Dynamic Input Selection: The framework dynamically selects inputs based on semantic diversity, predictive uncertainty, and potential judge biases, thereby more effectively capturing judge patterns and improving query efficiency.
  • Score Smoothing and Multi-Protocol Review: By applying score smoothing techniques, JudgeStealer preserves the ordinal structure of scores and mitigates catastrophic forgetting during surrogate adaptation through multi-protocol review.

Experimental Results

In extensive experiments on state-of-the-art LLM-as-a-judge and reward models, JudgeStealer consistently outperforms existing methods, achieving accuracies of 73.3%, 87.0%, and 71.6% for pointwise, pairwise, and listwise evaluation, respectively. Furthermore, JudgeStealer remains effective across different surrogate model scales, adaptation strategies, and reasoning settings, demonstrating robustness against representative extraction defenses.

Industry Impact and Developer Recommendations

  • Impact on AI Safety: The introduction of JudgeStealer highlights the potential security risks faced by LLM judging capabilities, underscoring the importance of strengthening model protection.
  • Recommendations for Developers: Developers should be aware of the threats posed by model extraction attacks and consider implementing stricter security measures, such as input validation and query limiting, in critical applications.
  • Future Research Directions: Further research could explore more sophisticated methods for leveraging cross-protocol agreement and develop customized extraction defense strategies for specific domains.

Conclusion

JudgeStealer provides an efficient and robust technical solution for replicating LLM judging capabilities, demonstrating the potential of extracting model capabilities in multi-protocol environments. This framework not only advances the field of AI safety but also lays the groundwork for future research in model extraction.


Source: ArXiv cs.CL (2026-08-27)

— END —

Tags: #JudgeStealer #Model Extraction #Large Language Models #AI Safety #Evaluation Protocols

Community Comments

Loading live comments and annotations…