CORL Team Releases NAMVIS: Revolutionizing Sparse-View Multi-View Image Synthesis
By Mr.Xu
Published:
Summary:The CORL Team introduces NAMVIS, a novel framework for sparse-view multi-view image synthesis. NAMVIS employs a geometry-conditioned next-scale autoregressive approach, replacing the traditional iterative denoising process of diffusion-based methods for more efficient multi-view generation. The framework outperforms existing diffusion baselines across multiple benchmarks, including PSNR, SSIM, and LPIPS, while achieving over a 3x speedup in inference time. The release of NAMVIS marks a significa
Background and Challenges
Sparse-view novel view synthesis is a central problem in 3D content creation. Traditional diffusion-based methods, while capable of generating high-quality images, suffer from high computational costs due to their iterative denoising process, particularly during inference, which limits their practical application.
Core Innovations of NAMVIS
NAMVIS addresses these challenges through the following innovations:
-
Geometry-Conditioned Next-Scale Autoregression: NAMVIS reformulates multi-view image synthesis as a geometry-conditioned next-scale autoregressive process, predicting discrete visual tokens through a small number of coarse-to-fine scale steps and sampling all tokens within each scale and across target views in parallel.
-
Multi-scale Projective Pose Encoding: This technique injects source and target camera transformations into both target-view self-attention and source-to-target cross-attention mechanisms, anchoring the generation process to explicit camera geometry.
-
Global Conditioning with Dense Geometry-Aware Cross-Attention: NAMVIS combines global conditioning with dense geometry-aware cross-attention, enabling the model to preserve source-view appearance while maintaining target-view consistency.
Experimental Results and Performance
Experiments on datasets such as Objaverse, GSO, and OmniObject3D demonstrate that NAMVIS outperforms existing diffusion baselines in terms of PSNR, SSIM, and LPIPS. Additionally, NAMVIS achieves over a 3x speedup in inference time compared to the evaluated diffusion baselines under the same evaluation setting.
Technical Highlights
- Efficiency: Significantly improves inference speed through autoregressive methods.
- Multi-View Consistency: Maintains multi-view consistency using geometry conditions and cross-attention mechanisms.
- Flexibility: Applicable to a variety of 3D datasets, demonstrating broad applicability.
Industry Impact and Developer Recommendations
The release of NAMVIS opens new possibilities for 3D content creation, virtual reality, and augmented reality. Developers can leverage NAMVIS to generate high-quality multi-view images more efficiently, thereby accelerating the 3D content creation workflow. Furthermore, the architecture of NAMVIS provides new directions for future research, particularly in multimodal generation and geometry-aware modeling.
Future Outlook
With further optimization and expansion, NAMVIS is expected to achieve more efficient and accurate multi-view image synthesis in more application scenarios, driving continuous advancements in 3D content creation technology.
— END —Source: Hugging Face Daily Papers (2026-10-03)
Tags: #NAMVIS #Multi-View Synthesis #Autoregressive Models #3D Content Creation #Geometry-Aware
Community Comments