Hugging Face Introduces NowWAM: Revolutionizing Generative Adaptation for Robot Control
By Mr.Xu
Published:
Summary:Hugging Face's research team introduces NowWAM, a novel framework for robot control that denoises current observations and predicts robot actions from the same visual stream, directly coupling generative adaptation with action prediction. Unlike traditional methods that rely on future visual targets, NowWAM achieves superior performance on the LIBERO-Plus benchmark, improving control accuracy and training efficiency while reducing the number of visual tokens and inference time, marking a signifi
1. Background and Motivation
Pretrained generative Diffusion Transformers (DiTs) have demonstrated remarkable capabilities in capturing rich pixel-level visual and language-conditioned structures through large-scale image and video generation training. However, the question of how to transfer these generative priors to robot control remains unresolved. Existing approaches typically rely on future visual prediction, but this method has several limitations.
2. Introduction to NowWAM Framework
NowWAM is a future-target-free co-training formulation that denoises the current observation and predicts robot actions from the same visual stream, directly coupling the generative objective with action prediction. The key innovations of this approach include:
- Denoising Trajectory as Control Interface: Utilizes the denoising trajectory as an interface for action prediction, eliminating the need for future visual targets.
- Visual Stream Coupling: Couples the generative objective with action prediction, enhancing the model's adaptability and robustness.
- Efficient Training: Reduces the number of training visual tokens and inference time, significantly improving training efficiency.
3. Experimental Results and Performance
On the LIBERO-Plus benchmark, NowWAM, when combined with FLUX2-Klein, achieved a control accuracy of 87.7%, outperforming the future-target co-training baseline by 6.1 percentage points. It also halved the number of training visual tokens (from 784 to 392) and reduced inference time from 2.85 seconds to 1.63 seconds, achieving a 1.8x speedup. Furthermore, with the pure text-to-image Z-Image backbone, NowWAM further reached 87.8% control accuracy, demonstrating that strong control adaptation is not tied to video generation or image-editing backbones.
4. Technical Highlights and Advantages
- No Need for Future Visual Targets: Simplifies the training process and reduces the demand for complex data.
- Efficient Training and Inference: Improves overall efficiency by reducing the number of visual tokens and inference time.
- Cross-Scenario Adaptability: Demonstrates strong generalization capabilities in both in-distribution and out-of-distribution scenarios.
5. Industry Impact and Developer Recommendations
NowWAM offers a new paradigm for robot control, particularly in complex environments. Developers are advised to:
- Experiment with applying NowWAM to different robot platforms to evaluate its performance in various scenarios.
- Combine NowWAM with other advanced AI technologies, such as reinforcement learning and imitation learning, to further enhance control performance.
- Stay updated with Hugging Face's latest research to gain more insights into generative adaptation and robot control.
6. Conclusion
NowWAM showcases the significant potential of generative adaptation in robot control, paving the way for future research and technological development.
— END —Source: Hugging Face Daily Papers (2026-09-23)
Tags: #Hugging Face #Robot Control #Generative Adaptation #NowWAM #DiTs
Community Comments