Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model
2023/10/23 by Shi, Ruoxi, Chen, Hansheng, Zhang, Zhuoyang +6 · 134 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Graphics (cs.GR)
paper · doi:10.48550/arxiv.2310.15110
Abstract
We report Zero123++, an image-conditioned diffusion model for generating 3D-consistent multi-view images from a single input view. To take full advantage of pretrained 2D generative priors, we develop various conditioning and training schemes to minimize the effort of finetuning from off-the-shelf image diffusion models such as Stable Diffusion. Zero123++ excels in producing high-quality, consistent multi-view images from a single image, overcoming common issues like texture degradation and geometric misalignment. Furthermore, we showcase the feasibility of training a ControlNet on Zero123++ for enhanced control over the generation process. The code is available at https://github.com/SUDO-AI-3D/zero123plus.
Cited by
- MVGBench: Comprehensive Benchmark for Multi-view Generation Models
- Fashion-3DLR: A Controllable 3D Garment Generation Using Pairwise Fashion Elements for Intelligent Design
- UMAMI: Unifying Masked Autoregressive Models and Deterministic Rendering for View Synthesis
- Learning High-Quality Initial Noise for Single-View Synthesis with Diffusion Models
- ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models
- AnchorHOI: Zero-shot Generation of 4D Human-Object Interaction via Anchor-based Prior Distillation
- CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
- FactorPortrait: Controllable Portrait Animation via Disentangled Expression, Pose, and Viewpoint
- OmniView: An All-Seeing Diffusion Model for 3D and 4D View Synthesis
- MeshRipple: Structured Autoregressive Generation of Artist-Meshes
- CAMEO: Correspondence-Attention Alignment for Multi-View Diffusion Models
- Efficient and Scalable Monocular Human-Object Interaction Motion Reconstruction
- PAT3D: Physics-Augmented Text-to-3D Scene Generation
- Single Image to High-Quality 3D Object via Latent Features
- Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approach
- ArticFlow: Generative Simulation of Articulated Mechanisms
- SVG360: Editable Multiview Vector Graphics from a Single SVG
- NaTex: Seamless Texture Generation as Latent Color Diffusion
- LSS3D: Learnable Spatial Shifting for Consistent and High-Quality 3D Generation from Single-Image
- GeoMVD: Geometry-Enhanced Multi-View Generation Model Based on Geometric Information Extraction
- 4DSTR: Advancing Generative 4D Gaussians with Spatial-Temporal Rectification for High-Quality and Consistent 4D Generation
- MagicView: Multi-View Consistent Identity Customization via Priors-Guided In-Context Learning
- FreeArt3D: Training-Free Articulated Object Generation using 3D Diffusion
- TRELLISWorld: Training-Free World Generation from Object Generators
- Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture Infilling
- GeoDiffusion: A Training-Free Framework for Accurate 3D Geometric Conditioning in Image Generation
- CUPID: Generative 3D Reconstruction via Joint Object and Pose Modeling
- Advances in 4D Representation: Geometry, Motion, and Interaction
- MV-Performer: Taming Video Diffusion Model for Faithful and Synchronized Multi-view Performer Synthesis
- Generating Surface for Text-to-3D using 2D Gaussian Splatting
- Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models
- VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
- Scaling Sequence-to-Sequence Generative Neural Rendering
- UP2You: Fast Reconstruction of Yourself from Unconstrained Photo Collections
- Drag4D: Align Your Motion with Text-Driven 3D Scene Generation
- Large Material Gaussian Model for Relightable 3D Generation
- VideoFrom3D: 3D Scene Video Generation via Complementary Image and Video Diffusion Models
- MMPart: Harnessing Multi-Modal Large Language Models for Part-Aware 3D Generation
- Dream3DAvatar: Text-Controlled 3D Avatar Reconstruction from a Single Image
- HoloGarment: 360° Novel View Synthesis of In-the-Wild Garments
- Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation
- T2Bs: Text-to-Character Blendshapes via Video Generation
- DreamLifting: A Plug-in Module Lifting MV Diffusion Models for 3D Asset Generation
- SynthDrive: Scalable Real2Sim2Real Sensor Simulation Pipeline for High-Fidelity Asset Generation and Driving Data Synthesis
- Learning in ImaginationLand: Omnidirectional Policies through 3D Generative Models (OP-Gen)
- MV-RAG: Retrieval Augmented Multiview Diffusion
- UniView: Enhancing Novel View Synthesis From A Single Image By Unifying Reference Features
- Hyper Diffusion Avatars: Dynamic Human Avatar Generation using Network Weight Space Diffusion
- Unifi3D: A Study on 3D Representations for Generation and Reconstruction in a Common Framework
- Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion
- Droplet3D: Commonsense Priors from Videos Facilitate 3D Generation
- DATR: Diffusion-based 3D Apple Tree Reconstruction Framework with Sparse-View
- ObjFiller-3D: Consistent Multi-view 3D Inpainting via Video Diffusion Models
- Novel View Synthesis using DDIM Inversion
- A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation
- Make Your MoVe: Make Your 3D Contents by Adapting Multi-View Diffusion Models to External Editing
- CharacterShot: Controllable and Consistent 4D Character Animation
- CLONE: Continuous Latent Optimization for Normal Estimation via 3D Gaussian Splatting
- DualMat: PBR Material Estimation via Coherent Dual-Path Diffusion
- EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion
- Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis
- XSpecMesh: Quality-Preserving Auto-Regressive Mesh Generation Acceleration via Multi-Head Speculative Decoding
- BANG: Dividing 3D Assets via Generative Exploded Dynamics
- MVG4D: Image Matrix-Based Multi-View and Motion Generation for 4D Content Creation from a Single Image
- Stereo-GS: Multi-View Stereo Vision Model for Generalizable 3D Gaussian Splatting Reconstruction
- Towards Geometric and Textural Consistency 3D Scene Generation via Single Image-guided Model Generation and Layout Optimization
- Advances in Feed-Forward 3D Reconstruction and View Synthesis: A Survey
- SmokeSVD: Smoke Reconstruction from A Single View via Progressive Novel View Synthesis and Refinement with Diffusion Models
- Reconstruct, Inpaint, Test-Time Finetune: Dynamic Novel-view Synthesis from Monocular Videos
- Elevating 3D Models: High-Quality Texture and Geometry Refinement from a Low-Quality Model
- Orientation Matters: Making 3D Generative Models Orientation-Aligned
- InstaScene: Towards Complete 3D Instance Decomposition and Reconstruction from Cluttered Scenes
- Stable-Hair v2: Real-World Hair Transfer via Multiple-View Diffusion Model
- EscherNet++: Simultaneous Amodal Completion and Scalable View Synthesis through Masked Fine-Tuning and Enhanced Feed-Forward 3D Reconstruction
- DreamArt: Generating Interactable Articulated Objects from a Single Image
- DreamGrasp: Zero-Shot 3D Multi-Object Reconstruction from Partial-View Images for Robotic Manipulation
- OmniPart: Part-Aware 3D Generation with Semantic Decoupling and Structural Cohesion
- DreamComposer++: Empowering Diffusion Models with Multi-View Conditions for 3D Content Generation
- WAVE: Warp-Based View Guidance for Consistent Novel View Synthesis Using a Single Image
- UA-Pose: Uncertainty-Aware 6D Object Pose Estimation and Online Object Completion with Partial References
- Shape-for-Motion: Precise and Consistent Video Editing with 3D Proxy
- DeOcc-1-to-3: 3D De-Occlusion from a Single Image via Self-Supervised Multi-View Diffusion
- DreamAnywhere: Object-Centric Panoramic 3D Scene Generation
- EditP23: 3D Editing via Propagation of Image Prompts to Multi-View
- VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
- Auto-Regressively Generating Multi-View Consistent Images
- DualTHOR: A Dual-Arm Humanoid Simulation Platform for Contingency-Aware Planning
- Hunyuan3D 2.5: Towards High-Fidelity 3D Assets Generation with Ultimate Details
- Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material
- ProSplat: Improved Feed-Forward 3D Gaussian Splatting for Wide-Baseline Sparse Views
- WildCAT3D: Appearance-Aware Multi-View Diffusion in the Wild
- AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making
- SceneCompleter: Dense 3D Scene Completion for Generative Novel View Synthesis
- InstaInpaint: Instant 3D-Scene Inpainting with Masked Large Reconstruction Model
- AnimateAnyMesh: A Feed-Forward 4D Foundation Model for Text-Driven Universal Mesh Animation
- DreamCS: Geometry-Aware Text-to-3D Generation with Unpaired 3D Reward Supervision
- MeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data Augmentation
- OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
- FlexPainter: Flexible and Multi-View Consistent Texture Generation
- ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
- PromptVFX: Text-Driven Fields for Open-World 3D Gaussian Animation
- Pro3D-Editor : A Progressive-Views Perspective for Consistent and Precise 3D Editing
- AdaHuman: Animatable Detailed 3D Human Generation with Compositional Multiview Diffusion
- InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance Parsing
- Zero-P-to-3: Zero-Shot Partial-View Images to 3D Object
- PacTure: Efficient PBR Texture Generation on Packed Views with Visual Autoregressive Models
- Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction
- Styl3R: Instant 3D Stylized Reconstruction for Arbitrary Scenes and Styles
- ART-DECO: Arbitrary Text Guidance for 3D Detailizer Construction
- Refining Few-Step Text-to-Multiview Diffusion via Reinforcement Learning
- Semantic Compression of 3D Objects for Open and Collaborative Virtual Worlds
- MVPainter: Accurate and Detailed 3D Texture Generation via Multi-View Diffusion with Geometric Control
- ACT-R: Adaptive Camera Trajectories for Single View 3D Reconstruction
- CMD: Controllable Multiview Diffusion for 3D Editing and Progressive Generation
- Anymate: A Dataset and Baselines for Learning 3D Object Rigging
- DiffLocks: Generating 3D Hair from a Single Image using Diffusion Models
- Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
- 3D Stylization via Large Reconstruction Model
- Stream3D: Sequential Multi-View 3D Generation via Evidential Memory
- Pose-Aware Diffusion for 3D Generation
- TreeON: Reconstructing 3D Tree Point Clouds from Orthophotos and Heightmaps
- Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation
- LaRI: Layered Ray Intersections for Single-view 3D Geometric Reasoning
- DiMeR: Disentangled Mesh Reconstruction Model
- Any3DAvatar: Fast and High-Quality Full-Head 3D Avatar Reconstruction from Single Portrait Image
- TwoSquared: 4D Generation from 2D Image Pairs
- SOPHY: Learning to Generate Simulation-Ready Objects with Physical Materials
- HiScene: Creating Hierarchical 3D Scenes with Isometric View Generation
- VideoPanda: Video Panoramic Diffusion with Multi-view Attention
- SpinMeRound: Consistent Multi-View Identity Generation Using Diffusion Models
- I Want It That Way! Specifying Nuanced Camera Motions in Video Editing
- Text To 3D Object Generation For Scalable Room Assembly
- In-2-4D: Inbetweening from Two Single-View Images to 4D Generation
- Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D Objects
Related