LIRM: Large Inverse Rendering Model for Progressive Reconstruction of Shape, Materials and View-dependent Radiance Fields
2025/04/28 by Zhengqin Li, Li, Zhengqin, Dilin Wang +25 · 5 citations
Engineering · Computer Science · #3D Shape Modeling and Analysis #Computer Graphics and Visualization Techniques #Advanced Vision and Imaging
paper · pdf · doi:10.48550/arxiv.2504.20026
Abstract
We present Large Inverse Rendering Model (LIRM), a transformer architecture that jointly reconstructs high-quality shape, materials, and radiance fields with view-dependent effects in less than a second. Our model builds upon the recent Large Reconstruction Models (LRMs) that achieve state-of-the-art sparse-view reconstruction quality. However, existing LRMs struggle to reconstruct unseen parts accurately and cannot recover glossy appearance or generate relightable 3D contents that can be consumed by standard Graphics engines. To address these limitations, we make three key technical contributions to build a more practical multi-view 3D reconstruction framework. First, we introduce an update model that allows us to progressively add more input views to improve our reconstruction. Second, we propose a hexa-plane neural SDF representation to better recover detailed textures, geometry and material parameters. Third, we develop a novel neural directional-embedding mechanism to handle view-dependent effects. Trained on a large-scale shape and material dataset with a tailored coarse-to-fine training scheme, our model achieves compelling results. It compares favorably to optimization-based dense-view inverse rendering methods in terms of geometry and relighting accuracy, while requiring only a fraction of the inference time.
Citations
- Digital Twin Catalog: A Large-Scale Photorealistic 3D Object Digital Twin Dataset
- RelitLRM: Generative Relightable Radiance for Large Reconstruction Models
- SpaRP: Fast 3D Object Reconstruction and Pose Estimation from Sparse Views
- SF3D: Stable Fast 3D Mesh Reconstruction with UV-unwrapping and Illumination Disentanglement
- Meta 3D AssetGen: Text-to-Mesh Generation with High-Quality Geometry, Texture, and PBR Materials
- Neural Directional Encoding for Efficient and Accurate View-Dependent Appearance Modeling
- GS-LRM: Large Reconstruction Model for 3D Gaussian Splatting
- DreamCraft: Text-Guided Generation of Functional 3D Environments in Minecraft
- MeshLRM: Large Reconstruction Model for High-Quality Meshes
- InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models
- TripoSR: Fast 3D Object Reconstruction from a Single Image
- MVDiffusion++: A Dense High-resolution Multi-view Diffusion Model for Single or Sparse-view 3D Object Reconstruction
- SHINOBI: Shape and Illumination using Neural Object Decomposition via BRDF Optimization In-the-wild
- Objects With Lighting: A Real-World Dataset for Evaluating Reconstruction and Rendering for Object Relighting
- SteinDreamer: Variance Reduction for Text-to-3D Score Distillation via Stein Identity
- Taming Mode Collapse in Score Distillation for Text-to-3D Generation
- ZeroRF: Fast Sparse View 360° Reconstruction with Zero Pretraining
- GIR: 3D Gaussian Inverse Rendering for Relightable Scene Factorization
- HyperDreamer: Hyper-Realistic 3D Content Generation and Editing from a Single Image
- GaussianShader: 3D Gaussian Splatting with Shading Functions for Reflective Surfaces
- GS-IR: 3D Gaussian Splatting for Inverse Rendering
- PF-LRM: Pose-Free Large Reconstruction Model for Joint Pose and Shape Prediction
- DMV3D: Denoising Multi-View Diffusion using 3D Large Reconstruction Model
- One-2-3-45++: Fast Single Image to 3D Objects with Consistent Multi-View Generation and 3D Diffusion
- Instant3D: Fast Text-to-3D with Sparse-View Generation and Large Reconstruction Model
- LRM: Large Reconstruction Model for Single Image to 3D
- Stanford-ORB: A Real-World 3D Object Inverse Rendering Benchmark
- Wonder3D: Single Image to 3D using Cross-Domain Diffusion
- SIRe-IR: Inverse Rendering for BRDF Reconstruction with Shadow and Illumination Removal in High-Illuminance Scenes
- SyncDreamer: Generating Multiview-consistent Images from a Single-view Image
- MVDream: Multi-view Diffusion for 3D Generation
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- Objaverse-XL: A Universe of 10M+ 3D Objects
- Neuralangelo: High-Fidelity Neural Surface Reconstruction
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation
- Neural-PBIR Reconstruction of Shape, Material, and Illumination
- TensoIR: Tensorial Inverse Rendering
- NeMF: Inverse Volume Rendering with Neural Microflake Field
- Zero-1-to-3: Zero-shot One Image to 3D Object
- MVImgNet: A Large-scale Dataset of Multi-view Images
- K-Planes: Explicit Radiance Fields in Space, Time, and Appearance
- Scalable Diffusion Models with Transformers
- Objaverse: A Universe of Annotated 3D Objects
- Magic3D: High-Resolution Text-to-3D Content Creation
- NerfAcc: A General NeRF Acceleration Toolbox
- DreamFusion: Text-to-3D using 2D Diffusion
- ARF: Artistic Radiance Fields
- SparseNeuS: Fast Generalizable Neural Surface Reconstruction from Sparse Views
- Shape, Light, and Material Decomposition from Images using Monte Carlo Rendering and Denoising
- SAMURAI: Shape And Material from Unconstrained Real-world Arbitrary Image collections
- Physically-Based Editing of Indoor Scene Lighting from a Single Image
- Modeling Indirect Illumination for Inverse Rendering
- IRON: Inverse Rendering by Optimizing Neural SDFs and Materials from Photometric Images
- Instant Neural Graphics Primitives with a Multiresolution Hash Encoding
- InfoNeRF: Ray Entropy Minimization for Few-Shot Neural Volume Rendering
- RegNeRF: Regularizing Neural Radiance Fields for View Synthesis from Sparse Inputs
- GeoNeRF: Generalizing NeRF with Geometry Priors
- Extracting Triangular 3D Models, Materials, and Lighting From Images
- Neural-PIL: Neural Pre-Integrated Lighting for Reflectance Decomposition
- ABO: Dataset and Benchmarks for Real-World 3D Object Understanding
- Learning Indoor Inverse Rendering with 3D Spatially-Varying Lighting
- Volume Rendering of Neural Implicit Surfaces
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction
- Emerging Properties in Self-Supervised Vision Transformers
- PhySG: Inverse Rendering with Spherical Gaussians for Physics-based Material Editing and Relighting
- Putting NeRF on a Diet: Semantically Consistent Few-Shot View Synthesis
- MVSNeRF: Fast Generalizable Radiance Field Reconstruction from Multi-View Stereo
- PlenOctrees for Real-time Rendering of Neural Radiance Fields
- IBRNet: Learning Multi-View Image-Based Rendering
- NeRD: Neural Reflectance Decomposition from Image Collections
- pixelNeRF: Neural Radiance Fields from One or Few Images
- Score-Based Generative Modeling through Stochastic Differential Equations
- Denoising Diffusion Probabilistic Models
- Lighthouse: Predicting Lighting Volumes for Spatially-Coherent Illumination
- Kaolin: A PyTorch Library for Accelerating 3D Deep Learning Research
- Inverse Rendering for Complex Indoor Scenes: Shape, Spatially-Varying Lighting and SVBRDF from a Single Image
- Soft Rasterizer: A Differentiable Renderer for Image-based 3D Reasoning
- CGIntrinsics: Better Intrinsic Image Decomposition through Physically-Based Rendering
- Materials for Masses: SVBRDF Acquisition with a Single Mobile Phone Image
- The Unreasonable Effectiveness of Deep Features as a Perceptual Metric
- LIME: Live Intrinsic Material Estimation
- Learning to Predict Indoor Illumination from a Single Image
- ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single Image
Cited by
Related