WorldSculpt: Compositional 3D World Generation from Grounded Videos
By Mr.Xu
Published: · 4 views
Summary:WorldSculpt is a novel 3D scene generation method that can create compositional 3D representations of cluttered scenes containing hundreds of objects from multi-view videos. By extending the Pixal3D model with a multi-view conditioning pathway, it grounds object generation in multiple posed observations, enabling the generation of individual object meshes in a shared world frame. The study also introduces UE-MeshyScene, a photorealistic benchmark for densely cluttered scenes. Experiments demonst
Background and Challenges
Generating 3D representations of complex scenes is a critical challenge in applications like gaming, AR/VR, simulation, and robotics. Traditional geometry-based approaches reconstruct scenes as single representations, leaving incomplete geometry in occluded regions. Existing compositional methods with generative priors are largely limited to relatively simple scenes due to their reliance on generative priors.
Methodology and Innovations
WorldSculpt addresses these challenges through the following:
- Multi-view Conditioning Pathway: Extends the Pixal3D model with a multi-view conditioning pathway, grounding object generation in multiple posed observations.
- Generalization from Single Objects to Complex Scenes: The model is fine-tuned entirely on single objects in canonical space but can generalize to large scenes with severe occlusion without any scene-level training.
- UE-MeshyScene Benchmark: Introduces a photorealistic benchmark of densely cluttered scenes with hundreds of objects, per-object annotations, and ground-truth meshes for evaluating model performance.
Experimental Results
Across single-object, controlled multi-object, and UE-MeshyScene evaluations, WorldSculpt consistently outperforms prior approaches, with larger gains as scene complexity and occlusion increase.
Applications and Extensions
The method is not limited to video generation and can convert generated 3DGS worlds, such as Marble and HY-World 2.0, into compositional mesh scenes, demonstrating its broad applicability.
Industry Impact and Developer Recommendations
- For Developers: This research provides a new approach to 3D scene generation, particularly in handling complex scenes and occlusion, showcasing its powerful capabilities. Developers can leverage this method to enhance the quality of scene generation in AR/VR, gaming, and robotics applications.
- For Industry: The technology is expected to advance 3D content creation tools and promote AI applications in simulation and robotics.
- Future Directions: Further optimization of the model to handle larger-scale scenes and exploration of its potential in dynamic scene generation.
— END —Source: Hugging Face Daily Papers (2026-09-04)
Tags: #3D Modeling #Computer Vision #Multi-View #Generative Models #Hugging Face
Community Comments