LRM: Large Reconstruction Model for Single Image to 3D
2023/11/08 by Yicong Hong, Kai Zhang, Hong, Yicong +17 · 138 citations
Computer Science · Engineering · #3D Shape Modeling and Analysis #Advanced Vision and Imaging #Artificial Intelligence (cs.AI) #Computer Graphics and Visualization Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Graphics (cs.GR) #Machine Learning (cs.LG)
paper · pdf · doi:10.48550/arxiv.2311.04400
openalex publication_date 2023/11/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We propose the first Large Reconstruction Model (LRM) that predicts the 3D model of an object from a single input image within just 5 seconds. In contrast to many previous methods that are trained on small-scale datasets such as ShapeNet in a category-specific fashion, LRM adopts a highly scalable transformer-based architecture with 500 million learnable parameters to directly predict a neural radiance field (NeRF) from the input image. We train our model in an end-to-end manner on massive multi-view data containing around 1 million objects, including both synthetic renderings from Objaverse and real captures from MVImgNet. This combination of a high-capacity model and large-scale training data empowers our model to be highly generalizable and produce high-quality 3D reconstructions from various testing inputs, including real-world in-the-wild captures and images created by generative models. Video demos and interactable 3D meshes can be found on our LRM project webpage: https://yiconghong.me/LRM.
Cited by
- Memorization in 3D Shape Generation: An Empirical Study
- ShapeR: Robust Conditional 3D Shape Generation from Casual Captures
- VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement
- RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
- Pixal3D: Pixel-Aligned 3D Generation from Images
- Feedforward 3D Editing Learns from Semantic-Part Transformation
- Geometric Context Transformer for Streaming 3D Reconstruction
- UltraShape 1.0: High-Fidelity 3D Shape Generation via Scalable Geometric Refinement
- FlexAvatar: Flexible Large Reconstruction Model for Animatable Gaussian Head Avatars with Detailed Deformation
- Make-It-Poseable: Feed-forward Latent Posing Model for 3D Characters
- ART: Articulated Reconstruction Transformer
- SS4D: Native 4D Generative Model via Structured Spacetime Latents
- An evaluation of SVBRDF Prediction from Generative Image Models for Appearance Modeling of 3D Scenes
- Feedforward 3D Editing via Text-Steerable Image-to-3D
- SplatPainter: Interactive Authoring of 3D Gaussians from 2D Edits via Test-Time Training
- Sharp Monocular View Synthesis in Less Than a Second
- Long-LRM++: Preserving Fine Details in Feed-Forward Wide-Coverage Reconstruction
- Photo3D: Advancing Photorealistic 3D Generation through Structure-Aligned Detail Enhancement
- Voxify3D: Pixel Art Meets Volumetric Rendering
- ViSA: 3D-Aware Video Shading for Real-Time Upper-Body Avatar Creation
- MoCA: Mixture-of-Components Attention for Scalable Compositional 3D Generation
- MeshRipple: Structured Autoregressive Generation of Artist-Meshes
- Tessellation GS: Neural Mesh Gaussians for Robust Monocular Reconstruction of Dynamic Objects
- Multi-view Pyramid Transformer: Look Coarser to See Broader
- DragMesh: Interactive 3D Generation Made Easy
- LaFiTe: A Generative Latent Field for 3D Native Texturing
- Flux4D: Flow-based Unsupervised 4D Reconstruction
- Controllable 3D Object Generation with Single Image Prompt
- Cue3D: Quantifying the Role of Image Cues in Single-Image 3D Generation
- Asset-Driven Sematic Reconstruction of Dynamic Scene with Multi-Human-Object Interactions
- CC-FMO: Camera-Conditioned Zero-Shot Single Image to 3D Scene Generation with Foundation Model Orchestration
- PAT3D: Physics-Augmented Text-to-3D Scene Generation
- Bringing Your Portrait to 3D Presence
- ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models
- AnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned Flows
- ShapeGen: Towards High-Quality 3D Shape Synthesis
- LATTICE: Democratize High-Fidelity 3D Generation at Scale
- Single Image to High-Quality 3D Object via Latent Features
- NeAR: Coupled Neural Asset-Renderer Stack
- Native 3D Editing with Full Attention
- SVG360: Editable Multiview Vector Graphics from a Single SVG
- TRIM: Scalable 3D Gaussian Diffusion Inference with Temporal and Spatial Trimming
- NaTex: Seamless Texture Generation as Latent Color Diffusion
- Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation
- LSS3D: Learnable Spatial Shifting for Consistent and High-Quality 3D Generation from Single-Image
- Appreciate the View: A Task-Aware Evaluation Framework for Novel View Synthesis
- LARM: A Large Articulated-Object Reconstruction Model
- ProcGen3D: Learning Neural Procedural Graph Representations for Image-to-3D Reconstruction
- YoNoSplat: You Only Need One Model for Feedforward 3D Gaussian Splatting
- Adaptive 3D Reconstruction via Diffusion Priors and Forward Curvature-Matching Likelihood Updates
- Faithful Contouring: Near-Lossless 3D Voxel Representation Free from Iso-surface
- Wonder3D++: Cross-domain Diffusion for High-fidelity 3D Generation from a Single Image
- HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D Segmentation
- FullPart: Generating each 3D Part at Full Resolution
- FreeArt3D: Training-Free Articulated Object Generation using 3D Diffusion
- PixARMesh: Autoregressive Mesh-Native Single-View Scene Reconstruction
- Surflo: Consistent 3D Surface Flow Model with Global State
- NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction
- TRELLISWorld: Training-Free World Generation from Object Generators
- Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture Infilling
- ReconViaGen: Towards Accurate Multi-view 3D Object Reconstruction via Generation
- GeoDiffusion: A Training-Free Framework for Accurate 3D Geometric Conditioning in Image Generation
- WorldGrow: Generating Infinite 3D World
- CUPID: Generative 3D Reconstruction via Joint Object and Pose Modeling
- OnlineSplatter: Pose-Free Online 3D Reconstruction for Free-Moving Objects
- Positional Encoding Field
- Advances in 4D Representation: Geometry, Motion, and Interaction
- TOUCH: Text-guided Controllable Generation of Free-Form Hand-Object Interactions
- Capture, Canonicalize, Splat: Zero-Shot 3D Gaussian Avatars from Unstructured Phone Images
- Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator
- SViM3D: Stable Video Material Diffusion for Single Image 3D Generation
- Scaling Sequence-to-Sequence Generative Neural Rendering
- HART: Human Aligned Reconstruction Transformer
- LVT: Large-Scale Scene Reconstruction via Local View Transformers
- RapidMV: Leveraging Spatio-Angular Representations for Efficient and Consistent Text-to-Multi-View Synthesis
- NeoWorld: Neural Simulation of Explorable Virtual Worlds via Progressive 3D Unfolding
- UniLat3D: Geometry-Appearance Unified Latents for Single-Stage 3D Generation
- Towards Fine-Grained Text-to-3D Quality Assessment: A Benchmark and A Two-Stage Rank-Learning Metric
- Large Material Gaussian Model for Relightable 3D Generation
- Hunyuan3D-Omni: A Unified Framework for Controllable Generation of 3D Assets
- FreeInsert: Personalized Object Insertion with Geometric and Style Control
- Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation
- SemanticGarment: Semantic-Controlled Generation and Editing of 3D Gaussian Garments
- HyPlaneHead: Rethinking Tri-plane-like Representations in Full-Head Image Synthesis
- MeshSplat: Generalizable Sparse-View Surface Reconstruction via Gaussian Splatting
- SPGen: Spherical Projection as Consistent and Flexible Representation for Single Image 3D Shape Generation
- Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation
- Align 3D Representation and Text Embedding for 3D Content Personalization
- DreamLifting: A Plug-in Module Lifting MV Diffusion Models for 3D Asset Generation
- Scaling Transformer-Based Novel View Synthesis Models with Token Disentanglement and Synthetic Data
- SynthDrive: Scalable Real2Sim2Real Sensor Simulation Pipeline for High-Fidelity Asset Generation and Driving Data Synthesis
- Few-step Flow for 3D Generation via Marginal-Data Transport Distillation
- MarkSplatter: Generalizable Watermarking for 3D Gaussian Splatting Model via Splatter Image Structure
- NeuralSVCD for Efficient Swept Volume Collision Detection
- 3D-LATTE: Latent Space 3D Editing from Textual Instructions
- Droplet3D: Commonsense Priors from Videos Facilitate 3D Generation
- FastAvatar: Towards Unified Fast High-Fidelity 3D Avatar Reconstruction with Large Gaussian Reconstruction Transformers
- DATR: Diffusion-based 3D Apple Tree Reconstruction Framework with Sparse-View
- VoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space
- FastMesh: Efficient Artistic Mesh Generation via Component Decoupling
- Follow My Hold: Hand-Object Interaction Reconstruction through Geometric Guidance
- ObjFiller-3D: Consistent Multi-view 3D Inpainting via Video Diffusion Models
- Collaborative Multi-Modal Coding for High-Quality 3D Generation
- PhysGM: Large Physical Gaussian Model for Feed-Forward 4D Synthesis
- VertexRegen: Mesh Generation with Continuous Level of Detail
- Make Your MoVe: Make Your 3D Contents by Adapting Multi-View Diffusion Models to External Editing
- Matrix-3D: Omnidirectional Explorable 3D World Generation
- CharacterShot: Controllable and Consistent 4D Character Animation
- RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
- GAP: Gaussianize Any Point Clouds with Text Guidance
- Hi3DEval: Advancing 3D Generation Evaluation with Hierarchical Validity
- MagicHOI: Leveraging 3D Priors for Accurate Hand-object Reconstruction from Short Monocular Video Clips
- DualMat: PBR Material Estimation via Coherent Dual-Path Diffusion
- 4DVD: Cascaded Dense-view Video Diffusion Model for High-quality 4D Content Generation
- OmniShape: Zero-Shot Multi-Hypothesis Shape and Pose Estimation in the Real World
- H3R: Hybrid Multi-view Correspondence for Generalizable 3D Reconstruction
- Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single Images
- EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion
- Can3Tok: Canonical 3D Tokenization and Latent Modeling of Scene-Level 3D Gaussians
- Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation
- Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis
- XSpecMesh: Quality-Preserving Auto-Regressive Mesh Generation Acceleration via Multi-Head Speculative Decoding
- iLRM: An Iterative Large 3D Reconstruction Model
- DepR: Depth Guided Single-view Scene Reconstruction with Instance-level Diffusion
- UFV-Splatter: Pose-Free Feed-Forward 3D Gaussian Splatting Adapted to Unfavorable Views
- HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels
- Towards Scalable Spatial Intelligence via 2D-to-3D Data Lifting
- MVG4D: Image Matrix-Based Multi-View and Motion Generation for 4D Content Creation from a Single Image
- Stereo-GS: Multi-View Stereo Vision Model for Generalizable 3D Gaussian Splatting Reconstruction
- Advances in Feed-Forward 3D Reconstruction and View Synthesis: A Survey
- Ultra3D: Efficient and High-Fidelity 3D Generation with Part Attention
- AutoPartGen: Autogressive 3D Part Generation and Discovery
- Reconstruct, Inpaint, Test-Time Finetune: Dynamic Novel-view Synthesis from Monocular Videos
- 3DGAA: Realistic and Robust 3D Gaussian-based Adversarial Attack for Autonomous Driving
- EgoAnimate: Generating Human Animations from Egocentric top-down Views
- From One to More: Contextual Part Latents for 3D Generation
- InstaScene: Towards Complete 3D Instance Decomposition and Reconstruction from Cluttered Scenes
- EscherNet++: Simultaneous Amodal Completion and Scalable View Synthesis through Masked Fine-Tuning and Enhanced Feed-Forward 3D Reconstruction
Related