ZICQ
中 Log in / Sign up
Newsroom Agentic #Robotics #Imitation Learning #Transformer #Causal Models #Sequence Modeling

WorldToken: A Novel Time-First Sequence Modeling Approach for Robotic Imitation Learning

Avatar of Mr.Xu

By Mr.Xu

Published: · 2 views

中文阅读 (Chinese) English Version

Summary:WorldToken introduces a novel time-first policy instantiation that fuses multiview images, proprioception, and task conditioning into a single 'world token' per policy timestep. A causal temporal Transformer models the resulting world-token sequence, while a diffusion action head generates action chunks. On 23 RoboCasa tasks, a policy with 85.3M parameters trained from scratch achieves a 59.45% mean closed-loop success rate using 2,900 generated demonstrations per task. The study demonstrates co


Key Technical Highlights

  1. Time-First Policy Instantiation: WorldToken integrates multiview images, proprioception, and task conditioning into a single 'world token' per timestep, enabling efficient modeling of complex robotic tasks.

  2. Causal Temporal Transformer: The method employs a causal temporal Transformer to model the resulting world-token sequence, capturing temporal dependencies.

  3. Diffusion Action Head: A diffusion action head generates action chunks, ensuring continuity and accuracy in action generation.

  4. Experimental Validation: On 23 RoboCasa tasks, a policy with 85.3M parameters achieves a 59.45% mean closed-loop success rate using 2,900 generated demonstrations per task. The study shows consistent gains with additional target-domain data and diminishing returns beyond moderate model sizes.

  5. Impact of History Length: Reducing visible history length significantly lowers closed-loop success, indicating the importance of temporal context for robotic task success.

Industry Implications and Developer Recommendations

  • A New Breakthrough in Robotics: WorldToken offers a novel approach to robotic imitation learning, particularly excelling in handling complex tasks and heterogeneous inputs.

  • Enhanced Data Efficiency: The method demonstrates the continuous performance improvement with additional target-domain data, suggesting that developers should focus on data quality and diversity during training.

  • Model Size Trade-offs: The study indicates diminishing returns beyond moderate model sizes, advising developers to balance performance and computational costs.

  • Importance of Temporal Context: Reducing visible history length impairs performance, highlighting the need for considering temporal context in practical applications.

Future Research Directions

  • Comparison with Other Sequence Modeling Methods: Further research is needed to compare WorldToken with other sequence modeling methods (e.g., RNNs, Transformers) to identify its areas of strength.

  • Application in Multi-Task Learning: Exploring the use of WorldToken in multi-task learning to assess its generalization capabilities and cross-task adaptability.

  • Hardware Acceleration and Optimization: Investigating the implementation and optimization of WorldToken on hardware accelerators to enhance its efficiency in practical applications.


Source: Hugging Face Daily Papers (2026-08-23)

— END —

Tags: #Robotics #Imitation Learning #Transformer #Causal Models #Sequence Modeling

Community Comments

Loading live comments and annotations…