Hugging Face Releases OSWorld-Pro: Revolutionizing Procedural Evaluation for Computer-Use Agents
By Mr.Xu
Published:
Summary:Hugging Face has launched OSWorld-Pro, a procedural evaluation framework for Computer-Use Agents (CUAs) that includes over 300 tasks and 2,800 subgoals, grounded in over 67,000 human annotations. Utilizing robust human-aligned LLM-Judges, OSWorld-Pro evaluates the fulfillment of its subgoals, revealing the progress and failure modes of models throughout a series of sequentially dependent tasks. This approach provides deeper insights compared to traditional end-state performance assessments, enab
Revolutionizing the Evaluation of Computer-Use Agents
Hugging Face has introduced OSWorld-Pro, a novel evaluation framework designed to address the limitations of existing methods for assessing Computer-Use Agents (CUAs). Traditional evaluation approaches primarily focus on the final output of agents at the end of tasks, often overlooking their performance and failure modes during task execution. OSWorld-Pro revolutionizes the evaluation process through the following features:
Key Features
- Procedural Evaluation: OSWorld-Pro encompasses over 300 tasks and 2,800 subgoals, enabling a detailed assessment of agents at each step of task execution.
- Large-Scale Human Annotations: The framework is grounded in over 67,000 human annotations, ensuring the accuracy and reliability of evaluations.
- LLM-Judge Evaluation: Utilizing LLM-Judges aligned with human judgment, OSWorld-Pro effectively identifies the progress and failure modes of agents throughout task completion.
Technical Highlights
- Task Decomposition and Subgoal Evaluation: By decomposing complex tasks into multiple subgoals, OSWorld-Pro provides a granular evaluation of agent performance, revealing specific issues at different stages.
- Human-Aligned Evaluation Criteria: The evaluation criteria of LLM-Judges are highly consistent with human judgment, ensuring the reliability and interpretability of assessment results.
Industry Impact
The release of OSWorld-Pro sets a new standard for CUA evaluation, enabling developers to more accurately identify shortcomings in agents' performance in complex tasks and make targeted improvements. This not only enhances the overall performance of agents but also facilitates the application and development of AI in various fields.
Recommendations for Developers
- Utilize OSWorld-Pro for Evaluation: Developers are encouraged to use OSWorld-Pro to evaluate existing CUAs, identify potential issues, and optimize agent performance.
- Focus on Procedural Improvements: Emphasize the performance of agents during task execution, not just the final outcome.
- Leverage Human Annotation Data: Use the large-scale human annotation data provided by OSWorld-Pro to further enhance the accuracy and reliability of evaluations.
— END —Source: Hugging Face Daily Papers (2026-09-21)
Tags: #Hugging Face #Intelligent Agents #Evaluation Framework #AI Evaluation #OSWorld-Pro
Community Comments