ZICQ
中 Log in / Sign up
Newsroom Agentic #DeskForge #GUI Understanding #Computer-Use Agents #Data Annotation #Vision-Language Models

DeskForge: A Controllable Desktop Environment for Dense Supervision to Enhance GUI Grounding in Computer-Use Agents

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:DeskForge, introduced by Said Gurbuz and colleagues, is a controllable desktop environment framework designed to generate large-scale, high-quality supervision data for computer-use agents. By composing real applications and simulating diverse desktop scenarios, DeskForge constructs DeskForge-1M, a dataset containing 1.2 million annotated desktop observations and 159.7 million element instances. Experiments demonstrate that vision-language models fine-tuned on DeskForge-1M show significant perfo


Core Breakthroughs

The core innovation of DeskForge lies in its controllability and ability to simulate real desktop environments, including:

  • Diverse Scenario Generation: By adjusting application states, content, window layouts, appearances, and resolutions, DeskForge generates highly diverse desktop scenarios.
  • Dense Annotations: It fuses screenshots, accessibility trees, and window geometry to produce richly annotated desktop observations.
  • Large-Scale Dataset Construction: DeskForge constructs DeskForge-1M, a dataset containing 1.2 million annotated desktop observations and 159.7 million element instances, providing ample supervision data for model training.

Technical Highlights

  1. Controllability: DeskForge allows developers to control various aspects of the desktop environment through parameterized settings, enabling the generation of tailored training data.
  2. Data Fusion: It integrates multiple data sources (e.g., screenshots, accessibility trees, and window geometry) to produce comprehensive annotated data.
  3. Performance Improvement: Models fine-tuned on DeskForge-1M show significant performance gains across multiple GUI understanding benchmarks. For example, Qwen3.5-4B's accuracy increased by 11.51 percentage points on ScreenSpot-Pro and 10.11 points on OSWorld-G.
  4. Long-Horizon Task Completion: DeskForge fine-tuned models also demonstrate significant progress in long-horizon task completion. For instance, Qwen3.5-4B's success rate in WebArena-Infinity and OpenApps tasks improved from 31/119 and 3/100 to 50/119 and 15/100, respectively.

Industry Impact

DeskForge provides a powerful tool for research in computer-use agents, particularly in GUI understanding and long-horizon task completion. Its controllability and large-scale data generation capabilities make it an invaluable resource for researchers and developers aiming to enhance model performance. Additionally, the open-source nature of DeskForge fosters collaboration and innovation within the AI community.

Developer Recommendations

  • Data Utilization: Researchers and developers are encouraged to utilize the DeskForge-1M dataset for model training and fine-tuning to improve GUI understanding capabilities.
  • Framework Extension: Developers can extend the DeskForge framework to generate more domain-specific desktop scenario data.
  • Performance Optimization: When applying DeskForge-fine-tuned models, it is recommended to further optimize them based on long-horizon task completion requirements.

Conclusion

DeskForge offers a novel perspective for research in computer-use agents, leveraging controllable desktop environment generation and large-scale annotated data to advance GUI understanding and long-horizon task completion capabilities.


Source: Hugging Face Daily Papers (2026-10-01)

— END —

Tags: #DeskForge #GUI Understanding #Computer-Use Agents #Data Annotation #Vision-Language Models

Community Comments

Loading live comments and annotations…