Objaverse-XL: A Universe of 10M+ 3D Objects
2023/07/11 by Matt Deitke, Deitke, Matt, Ruoshi Liu +31 · 230 citations
Computer Science · Earth and Planetary Sciences · #3D Surveying and Cultural Heritage #Advanced Neural Network Applications #Advanced Vision and Imaging #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · pdf · doi:10.48550/arxiv.2307.05663
openalex publication_date 2023/07/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Natural language processing and 2D vision models have attained remarkable proficiency on many tasks primarily by escalating the scale of training data. However, 3D vision tasks have not seen the same progress, in part due to the challenges of acquiring high-quality 3D data. In this work, we present Objaverse-XL, a dataset of over 10 million 3D objects. Our dataset comprises deduplicated 3D objects from a diverse set of sources, including manually designed objects, photogrammetry scans of landmarks and everyday items, and professional scans of historic and antique artifacts. Representing the largest scale and diversity in the realm of 3D datasets, Objaverse-XL enables significant new possibilities for 3D vision. Our experiments demonstrate the improvements enabled with the scale provided by Objaverse-XL. We show that by training Zero123 on novel view synthesis, utilizing over 100 million multi-view rendered images, we achieve strong zero-shot generalization abilities. We hope that releasing Objaverse-XL will enable further innovations in the field of 3D vision at scale.
Cited by
- MVGBench: Comprehensive Benchmark for Multi-view Generation Models
- Memorization in 3D Shape Generation: An Empirical Study
- ShapeR: Robust Conditional 3D Shape Generation from Casual Captures
- A Three-Level Alignment Framework for Large-Scale 3D Retrieval and Controlled 4D Generation
- Pixal3D: Pixel-Aligned 3D Generation from Images
- Orient Anything V2: Unifying Orientation and Rotation Understanding
- LiteGE: Lightweight Geodesic Embedding for Efficient Geodesics Computation and Non-Isometric Shape Correspondence
- MatLat: Material Latent Space for PBR Texture Generation
- Learning High-Quality Initial Noise for Single-View Synthesis with Diffusion Models
- Native and Compact Structured Latents for 3D Generation
- ART: Articulated Reconstruction Transformer
- SS4D: Native 4D Generative Model via Structured Spacetime Latents
- Particulate: Feed-Forward 3D Object Articulation
- ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models
- Feedforward 3D Editing via Text-Steerable Image-to-3D
- CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
- PoseGAM: Robust Unseen Object Pose Estimation via Geometry-Aware Multi-View Reasoning
- Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation
- MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular Videos
- UniPart: Part-Level 3D Generation with Unified 3D Geom-Seg Latents
- ASSIST-3D: Adapted Scene Synthesis for Class-Agnostic 3D Instance Segmentation
- Photo3D: Advancing Photorealistic 3D Generation through Structure-Aligned Detail Enhancement
- MeshRipple: Structured Autoregressive Generation of Artist-Meshes
- LaFiTe: A Generative Latent Field for 3D Native Texturing
- Not All Birds Look The Same: Identity-Preserving Generation For Birds
- TEXTRIX: Latent Attribute Grid for Native Texture Generation and Beyond
- TabletopGen: Tabletop Scene Generation and Interactive Simulation for Robotic Manipulation
- S2AM3D: Scale-controllable Part Segmentation of 3D Point Clouds
- Efficient and Scalable Monocular Human-Object Interaction Motion Reconstruction
- CC-FMO: Camera-Conditioned Zero-Shot Single Image to 3D Scene Generation with Foundation Model Orchestration
- Action-guided generation of 3D functionality segmentation data
- GOATex: Geometry & Occlusion-Aware Texturing
- ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models
- Wukong's 72 Transformations: High-fidelity Textured 3D Morphing via Flow Models
- AnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned Flows
- CaliTex: Geometry-Calibrated Attention for View-Coherent 3D Texture Generation
- Pygmalion Effect in Vision: Image-to-Clay Translation for Reflective Geometry Reconstruction
- MFM-point: Multi-scale Flow Matching for Point Cloud Generation
- VibraVerse: A Large-Scale Geometry-Acoustics Alignment Dataset for Physically-Consistent Multimodal Learning
- LumiTex: Towards High-Fidelity PBR Texture Generation with Illumination Context
- 4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models
- ShapeGen: Towards High-Quality 3D Shape Synthesis
- View-Consistent Diffusion Representations for 3D-Consistent Video Generation
- NI-Tex: Non-isometric Image-based Garment Texture Generation
- Neural Geometry Image-Based Representations with Optimal Transport (OT)
- RigAnyFace: Scaling Neural Facial Mesh Auto-Rigging with Unlabeled Data
- CrossJEPA: Cross-Modal Joint-Embedding Predictive Architecture for Efficient 3D Representation Learning from 2D Images
- MimiCAT: Mimic with Correspondence-Aware Cascade-Transformer for Category-Free 3D Pose Transfer
- ArticFlow: Generative Simulation of Articulated Mechanisms
- Native 3D Editing with Full Attention
- Free-Form Scene Editor: Enabling Multi-Round Object Manipulation like in a 3D Engine
- Wave-Former: Through-Occlusion 3D Reconstruction via Wireless Shape Completion
- 3DAlign-DAER: Dynamic Attention Policy and Efficient Retrieval Strategy for Fine-grained 3D-Text Alignment at Scale
- LSS3D: Learnable Spatial Shifting for Consistent and High-Quality 3D Generation from Single-Image
- LARM: A Large Articulated-Object Reconstruction Model
- Sat2RealCity: Geometry-Aware and Appearance-Controllable 3D Urban Generation from Satellite Imagery
- LoG3D: Ultra-High-Resolution 3D Shape Modeling via Local-to-Global Partitioning
- DensiCrafter: Physically-Constrained Generation and Fabrication of Self-Supporting Hollow Structures
- Faithful Contouring: Near-Lossless 3D Voxel Representation Free from Iso-surface
- LEGO-Eval: Towards Fine-Grained Evaluation on Synthesizing 3D Embodied Environments with Tool Augmentation
- DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos
- FullPart: Generating each 3D Part at Full Resolution
- FreeArt3D: Training-Free Articulated Object Generation using 3D Diffusion
- MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space
- Picasso: Holistic Scene Reconstruction with Physics-Constrained Sampling
- SynHLMA:Synthesizing Hand Language Manipulation for Articulated Object with Discrete Human Object Interaction Representation
- ReconViaGen: Towards Accurate Multi-view 3D Object Reconstruction via Generation
- WorldGrow: Generating Infinite 3D World
- CUPID: Generative 3D Reconstruction via Joint Object and Pose Modeling
- OnlineSplatter: Pose-Free Online 3D Reconstruction for Free-Moving Objects
- PartNeXt: A Next-Generation Dataset for Fine-Grained and Hierarchical 3D Part Understanding
- Advances in 4D Representation: Geometry, Motion, and Interaction
- GuideFlow3D: Optimization-Guided Rectified Flow For Appearance Transfer
- Procedural Scene Programs for Open-Universe Scene Generation: LLM-Free Error Correction via Program Search
- UrbanVerse: Scaling Urban Simulation by Watching City-Tour Videos
- Scene Coordinate Reconstruction Priors
- Jigsaw3D: Disentangled 3D Style Transfer via Patch Shuffling and Masking
- Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild
- Scaling Sequence-to-Sequence Generative Neural Rendering
- Towards Scalable and Consistent 3D Editing
- Text-to-Scene with Large Reasoning Models
- ASIA: Adaptive 3D Segmentation using Few Image Annotations
- Towards Fine-Grained Text-to-3D Quality Assessment: A Benchmark and A Two-Stage Rank-Learning Metric
- Sparse-Up: Learnable Sparse Upsampling for 3D Generation with High-Fidelity Textures
- ZeroScene: A Zero-Shot Framework for 3D Scene Generation from a Single Image and Controllable Texture Editing
- PartSAM: A Scalable Promptable Part Segmentation Model Trained on Native 3D Data
- QuadGPT: Native Quadrilateral Mesh Generation with Autoregressive Models
- SeamCrafter: Enhancing Mesh Seam Generation for Artist UV Unwrapping via Reinforcement Learning
- ArtUV: Artist-style UV Unwrapping
- PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation
- MeshMosaic: Scaling Artist Mesh Generation via Local-to-Global Assembly
- MeshFM: 2D Features Are All You Need for 3D Shape Understanding
- Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
- Hierarchical Neural Semantic Representation for 3D Semantic Correspondence
- AToken: A Unified Tokenizer for Vision
- CraftMesh: High-Fidelity Generative Mesh Manipulation via Poisson Seamless Fusion
- Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation
- T2Bs: Text-to-Character Blendshapes via Video Generation
- Structural Energy-Guided Sampling for View-Consistent Text-to-3D
- X-Part: high fidelity and structure coherent shape decomposition
- DreamLifting: A Plug-in Module Lifting MV Diffusion Models for 3D Asset Generation
- P3-SAM: Native 3D Part Segmentation
- SynthDrive: Scalable Real2Sim2Real Sensor Simulation Pipeline for High-Fidelity Asset Generation and Driving Data Synthesis
- Grasp-MPC: Closed-Loop Visual Grasping via Value-Guided Model Predictive Control
- Hyper Diffusion Avatars: Dynamic Human Avatar Generation using Network Weight Space Diffusion
- Weakly-Supervised Learning of Dense Functional Correspondences
- Unifi3D: A Study on 3D Representations for Generation and Reconstruction in a Common Framework
- NeuralSVCD for Efficient Swept Volume Collision Detection
- 3D-LATTE: Latent Space 3D Editing from Textual Instructions
- SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass
- Droplet3D: Commonsense Priors from Videos Facilitate 3D Generation
- LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding
- VoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space
- FastMesh: Efficient Artistic Mesh Generation via Component Decoupling
- Collaborative Multi-Modal Coding for High-Quality 3D Generation
- Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
- Puppeteer: Rig and Animate Your 3D Models
- TexVerse: A Universe of 3D Objects with High-Resolution Textures
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- Towards Spatially Consistent Image Generation: On Incorporating Intrinsic Scene Properties into Diffusion Models
- SHREC'25 Track on Multiple Relief Patterns: Report and Analysis
- VertexRegen: Mesh Generation with Continuous Level of Detail
- SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
- PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation
- Hi3DEval: Advancing 3D Generation Evaluation with Hierarchical Validity
- MagicHOI: Leveraging 3D Priors for Accurate Hand-object Reconstruction from Short Monocular Video Clips
- HOLODECK 2.0: Vision-Language-Guided 3D World Generation with Editing
- H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
- Toward Using Machine Learning as a Shape Quality Metric for Liver Point Cloud Generation
- DisCo3D: Distilling Multi-View Consistency for 3D Scene Editing
- MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
- DreamSat-2.0: Towards a General Single-View Asteroid 3D Reconstruction
- Hestia: Voxel-Face-Aware Hierarchical Next-Best-View Acquisition for Efficient 3D Reconstruction
- Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation
- Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis
- Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models
- iLRM: An Iterative Large 3D Reconstruction Model
- DepR: Depth Guided Single-view Scene Reconstruction with Instance-level Diffusion
- DISCOVERSE: Efficient Robot Simulation in Complex High-Fidelity Environments
- Advances in Feed-Forward 3D Reconstruction and View Synthesis: A Survey
- From Scan to Action: Leveraging Realistic Scans for Embodied Scene Understanding
- Ultra3D: Efficient and High-Fidelity 3D Generation with Part Attention
- Dens3R: A Foundation Model for 3D Geometry Prediction
- The Less You Depend, The More You Learn: Synthesizing Novel Views from Sparse, Unposed Images with Minimal 3D Knowledge
- PhysX-3D: Physical-Grounded 3D Asset Generation
- Elevating 3D Models: High-Quality Texture and Geometry Refinement from a Low-Quality Model
- Efficient Part-level 3D Object Generation via Dual Volume Packing
- Orientation Matters: Making 3D Generative Models Orientation-Aligned
- UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation
- Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation
- DreamArt: Generating Interactable Articulated Objects from a Single Image
- SV-DRR: High-Fidelity Novel View X-Ray Synthesis Using Diffusion Model
- LoomNet: Enhancing Multi-View Image Generation via Latent Space Weaving
- SeqTex: Generate Mesh Textures in Video Sequence
- MoReMouse: Monocular Reconstruction of Laboratory Mouse
- TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking
- WAVE: Warp-Based View Guidance for Consistent Novel View Synthesis Using a Single Image
- GIQ: Benchmarking 3D Geometric Reasoning of Vision Foundation Models with Simulated and Real Polyhedra
- AlignCVC: Aligning Cross-View Consistency for Single-Image-to-3D Generation
- 3D Shape Generation: A Survey
- Drive Any Mesh: 4D Latent Diffusion for Mesh Deformation from Video
- PhysRig: Differentiable Physics-Based Skinning and Rigging Framework for Realistic Articulated Object Modeling
- Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning
- EditP23: 3D Editing via Propagation of Image Prompts to Multi-View
- AnimaX: Animating the Inanimate in 3D with Joint Video-Pose Diffusion Models
- 4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time
- Light of Normals: Unified Feature Representation for Universal Photometric Stereo
- Auto-Regressive Surface Cutting
- PhysID: Physics-based Interactive Dynamics from a Single-view Image
- Assembler: Scalable 3D Part Assembly via Anchor Point Diffusion
- Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material
- ImmerseGen: Agent-Guided Immersive World Generation with Alpha-Textured Proxies
- WildCAT3D: Appearance-Aware Multi-View Diffusion in the Wild
- Edit360: 2D Image Edits to 3D Assets from Any Angle
- AnimateAnyMesh: A Feed-Forward 4D Foundation Model for Text-Driven Universal Mesh Animation
- You Only Estimate Once: Unified, One-stage, Real-Time Category-level Articulated Object 6D Pose Estimation for Robotic Grasping
- PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers
- Bridging Geometry-Coherent Text-to-3D Generation with Multi-View Diffusion Priors and Gaussian Splatting
- Seeing in the Dark: Benchmarking Egocentric 3D Vision with the Oxford Day-and-Night Dataset
- ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
- Pro3D-Editor : A Progressive-Views Perspective for Consistent and Precise 3D Editing
- ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image Intermediary
- Zero-P-to-3: Zero-Shot Partial-View Images to 3D Object
- PacTure: Efficient PBR Texture Generation on Packed Views with Visual Autoregressive Models
- MV-CoLight: Efficient Object Compositing with Consistent Lighting and Shadow Generation
- Is Single-View Mesh Reconstruction Ready for Robotics?
- Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention
- G-Skin: Learning to Bind 3D Gaussians with Generative Visual Priors
- Mesh-RFT: Enhancing Mesh Generation via Fine-grained Reinforcement Fine-Tuning
- Beyond Global Latents: Chunk-Based Sparse Grid VAE for Scalable 3D Modeling
- Constructing a 3D Scene from a Single Image
- X-GRM: Large Gaussian Reconstruction Model for Sparse-view X-rays to Computed Tomography
- Personalize Your Gaussian: Consistent 3D Scene Personalization from a Single Image
- Sparc3D: Sparse Representation and Construction for High-Resolution 3D Shapes Modeling
- FreeMesh: Boosting Mesh Generation with Coordinates Merging
- NeuSEditor: From Multi-View Images to Text-Guided Neural Surface Edits
- Procedural Generation of Articulated Simulation-Ready Assets
- AIMold: An Autonomous AI-based Pipeline for Complex Mold Design
- FoldNet: Learning Generalizable Closed-Loop Policy for Garment Folding via Keypoint-Driven Asset and Demonstration Synthesis
- Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets
- Text-to-CadQuery: A New Paradigm for CAD Generation with Scalable Large Model Capabilities
- RefRef: A Synthetic Dataset and Benchmark for Reconstructing Refractive and Reflective Objects
- Anymate: A Dataset and Baselines for Learning 3D Object Rigging
- Generating Physically Stable and Buildable Brick Structures from Text
- MVHumanNet++: A Large-scale Dataset of Multi-view Daily Dressing Human Captures with Richer Annotations for 3D Human Digitization
- Towards Autonomous Micromobility through Scalable Urban Simulation
- BiMotion: B-spline Motion for Text-guided Dynamic 3D Character Generation
- Syn4D: A Multiview Synthetic 4D Dataset
- Pose-Aware Diffusion for 3D Generation
- ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation
- CubePart: An Open-Vocabulary Part-Controllable 3D Generator
- LIRM: Large Inverse Rendering Model for Progressive Reconstruction of Shape, Materials and View-dependent Radiance Fields
- Boosting 3D Liver Shape Datasets with Diffusion Models and Implicit Neural Representations
- RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning
- Vysics: Object Reconstruction Under Occlusion by Fusing Vision and Contact-Rich Physics
- Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation
- Dynamic Camera Poses and Where to Find Them
- PICO: Reconstructing 3D People In Contact with Objects
- STP4D: Spatio-Temporal-Prompt Consistent Modeling for Text-to-4D Gaussian Splatting
- DiMeR: Disentangled Mesh Reconstruction Model
- StyleMe3D: Stylization with Disentangled Priors by Multiple Encoders on 3D Gaussians
- TwoSquared: 4D Generation from 2D Image Pairs
- One Model to Rig Them All: Diverse Skeleton Rigging with UniRig
- NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion Priors
- CAP-Net: A Unified Network for 6D Pose and Size Estimation of Categorical Articulated Parts from a Single RGB-D Image
- SpinMeRound: Consistent Multi-View Identity Generation Using Diffusion Models
- Digital Twin Catalog: A Large-Scale Photorealistic 3D Object Digital Twin Dataset
- BoxDreamer: Dreaming Box Corners for Generalizable Object Pose Estimation
- Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D Objects
- Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot Manipulation
- Imperative vs. Declarative Programming Paradigms for Open-Universe Scene Generation
Related