Reconstructing 4D Spatial Intelligence: A Survey
2025/07/28 by Cao, Yukang, Jiahao Lu, Lu, Jiahao +18 · 4 citations
Computer Science · Engineering · #3D Shape Modeling and Analysis #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robotics and Sensor-Based Localization
paper · pdf · doi:10.48550/arxiv.2507.21045
openalex publication_date 2025/07/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Reconstructing 4D spatial intelligence from visual observations has long been a central yet challenging task in computer vision, with broad real-world applications. These range from entertainment domains like movies, where the focus is often on reconstructing fundamental visual elements, to embodied AI, which emphasizes interaction modeling and physical realism. Fueled by rapid advances in 3D representations and deep learning architectures, the field has evolved quickly, outpacing the scope of previous surveys. Additionally, existing surveys rarely offer a comprehensive analysis of the hierarchical structure of 4D scene reconstruction. To address this gap, we present a new perspective that organizes existing methods into five progressive levels of 4D spatial intelligence: (1) Level 1 -- reconstruction of low-level 3D attributes (e.g., depth, pose, and point maps); (2) Level 2 -- reconstruction of 3D scene components (e.g., objects, humans, structures); (3) Level 3 -- reconstruction of 4D dynamic scenes; (4) Level 4 -- modeling of interactions among scene components; and (5) Level 5 -- incorporation of physical laws and constraints. We conclude the survey by discussing the key challenges at each level and highlighting promising directions for advancing toward even richer levels of 4D spatial intelligence. To track ongoing developments, we maintain an up-to-date project page: https://github.com/yukangcao/Awesome-4D-Spatial-Intelligence.
Citations
- π3: Permutation-Equivariant Visual Geometry Learning
- SpatialTrackerV2: 3D Point Tracking Made Easy
- Streaming 4D Visual Geometry Transformer
- Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer Memory
- Light of Normals: Unified Feature Representation for Universal Photometric Stereo
- SOF: Sorted Opacity Fields for Fast Unbounded Surface Reconstruction
- 3D Gaussian Splatting for Fine-Detailed Surface Reconstruction in Large-Scale Scene
- Multiview Geometric Regularization of Gaussian Splatting for Accurate Radiance Fields
- Multi-view Surface Reconstruction Using Normal and Reflectance Cues
- Photoreal Scene Reconstruction from an Egocentric Device
- UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation
- LODGE: Level-of-Detail Large-Scale Gaussian Splatting with Efficient Rendering
- PhysicsNeRF: Physics-Guided 3D Reconstruction from Sparse Views
- QuickSplat: Fast 3D Surface Reconstruction via Learned Gaussian Initialization
- Dynamic Camera Poses and Where to Find Them
- Seurat: From Moving Points to Depth
- TAPIP3D: Tracking Any Point in Persistent 3D Geometry
- Back on Track: Bundle Adjustment for Dynamic Scene Reconstruction
- ODHSR: Online Dense 3D Reconstruction of Humans and Scenes from Monocular Videos
- St4RTrack: Simultaneous 4D Reconstruction and Tracking in the World
- UniPhys: Unified Planner and Controller with Diffusion for Flexible Physics-Based Character Control
- Regist3R: Incremental Registration with Stereo Foundation Model
- The Invisible EgoHand: 3D Hand Forecasting through EgoBody Pose Estimation
- Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction
- POMATO: Marrying Pointmap Matching with Temporal Motion for Dynamic 3D Reconstruction
- D2USt3R: Enhancing 3D Reconstruction for Dynamic Scenes
- GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors
- Easi3R: Estimating Disentangled Motion from DUSt3R Without Training
- AnyCam: Learning to Recover Camera Poses and Intrinsics from Casual Videos
- CityGS-X: A Scalable Architecture for Efficient and Geometrically Accurate Large-Scale Scene Reconstruction
- DSO: Aligning 3D Generators with Simulation Feedback for Physical Soundness
- Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single Video
- GaussianUDF: Inferring Unsigned Distance Functions through 3D Gaussian Splatting
- Aether: Geometric-Aware Unified World Modeling
- Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene Priors
- Dynamic Point Maps: A Versatile Representation for Dynamic 3D Reconstruction
- VGGT: Visual Geometry Grounded Transformer
- MUSt3R: Multi-view Network for Stereo 3D Reconstruction
- CAST: Component-Aligned 3D Scene Reconstruction from an RGB Image
- HD-EPIC: A Highly-Detailed Egocentric Video Dataset
- ASAP: Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills
- Light3R-SfM: Towards Feed-forward Structure-from-Motion
- Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
- Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
- Continuous 3D Perception Model with Persistent State
- Zero-Shot Monocular Scene Flow Estimation in the Wild
- Joint Optimization for 4D Human-Scene Reconstruction in the Wild
- DAS3R: Dynamics-Aware Gaussian Splatting for Static Scene Reconstruction
- SolidGS: Consolidating Gaussian Surfel Splatting for Sparse-View Surface Reconstruction
- MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors
- Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos
- PBR-NeRF: Inverse Rendering with Physics-Based Neural Fields
- DGNS: Deformable Gaussian Splatting and Dynamic Neural Surface for Monocular Dynamic 3D Reconstruction
- MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos
- EgoPoints: Advancing Point Tracking for Egocentric Videos
- Align3R: Aligned Monocular Depth Estimation for Dynamic Videos
- Feed-Forward Bullet-Time Reconstruction of Dynamic Scenes from Monocular Videos
- Geometry Field Splatting with Gaussian Surfels
- Adaptive and Temporally Consistent Gaussian Surfels for Multi-view Dynamic Reconstruction
- Planar Reflection-Aware Neural Radiance Fields
- CityGaussianV2: Efficient and Geometrically Accurate Reconstruction for Large-Scale Scenes
- DELTA: Dense Efficient Long-range 3D Tracking for any video
- HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots
- Harmony4D: A Video Dataset for In-The-Wild Close Human Interactions
- SpectroMotion: Dynamic 3D Reconstruction of Specular Scenes
- DN-4DGS: Denoised Deformable Network with Temporal-Spatial Aggregation for Dynamic Scene Rendering
- Depth Any Video with Scalable Synthetic Data
- MotionGS: Exploring Explicit Motion Guidance for Deformable 3D Gaussian Splatting
- AvatarGO: Zero-shot 4D Human-Object Interaction Generation and Animation
- CLoSD: Closing the Loop between Simulation and Diffusion for multi-task character control
- MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
- Estimating Body and Hand Motion in an Ego-sensed World
- MASt3R-SfM: a Fully-Integrated Solution for Unconstrained Structure-from-Motion
- Space-time 2D Gaussian Splatting for Accurate Surface Reconstruction under Complex Dynamic Scenes
- EgoLM: Multi-Modal Language Model of Egocentric Motions
- MaskedMimic: Unified Physics-Based Character Control Through Masked Motion Inpainting
- HMD2: Environment-aware Motion Generation from Single Egocentric Head-Mounted Device
- AMEGO: Active Memory from long EGOcentric videos
- Surface-Centric Modeling for High-Fidelity Generalizable Neural Surface Reconstruction
- DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos
- Efficient Camera Exposure Control for Visual Odometry via Deep Reinforcement Learning
- 3D Reconstruction with Spatial Memory
- DynaSurfGS: Dynamic Surface Reconstruction with Planar-based Gaussian Splatting
- InterTrack: Tracking Human Object Interaction without Object Templates
- AvatarPose: Avatar-guided 3D Pose Estimation of Close Human Interaction from Sparse Multi-view Videos
- Deep Patch Visual SLAM
- 3D Gaussian Splatting: Survey, Technologies, Challenges, and Opportunities
- 3D Gaussian Splatting: Survey, Technologies, Challenges, and Opportunities
- SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency
- Shape of Motion: 4D Reconstruction from a Single Video
- Omnigrasp: Grasping Diverse Objects with Simulated Humanoids
- Neural Localizer Fields for Continuous 3D Human Pose and Shape Estimation
- Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
- CrowdMoGen: Zero-Shot Text-Driven Collective Motion Generation
- LaRa: Efficient Large-Baseline Radiance Fields
- Grounding Image Matching in 3D with MASt3R
- Nymeria: A Massive Collection of Multimodal Egocentric Daily Motion in the Wild
- Introducing HOT3D: An Egocentric Dataset for 3D Hand and Object Tracking
- Depth Anything V2
- A Survey on 3D Human Avatar Modeling -- From Reconstruction to Generation
- GenS: Generalizable Neural Surface Reconstruction from Multi-View Images
- Learning Temporally Consistent Video Depth from Video Diffusion Priors
- 4Diffusion: Multi-view Video Diffusion Model for 4D Generation
- Hierarchical World Models as Visual Whole-Body Humanoid Controllers
- MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds
- Splat-SLAM: Globally Optimized RGB-only SLAM with 3D Gaussians
- Diffusion4D: Fast Spatial-temporal Consistent 4D Generation via Video Diffusion Models
- PACER+: On-Demand Pedestrian Animation Controller in Driving Scenarios
- TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose Representation
- PhyRecon: Physically Plausible Neural Scene Reconstruction
- FlowMap: High-Quality Camera Poses, Intrinsics, and Depth via Gradient Descent
- PhysDreamer: Physics-Based Interaction with 3D Objects via Video Generation
- MultiPhys: Multi-Person Physics-aware 3D Motion Estimation
- REACTO: Reconstructing Articulated Objects from a Single Video
- Closely Interactive Human Reconstruction with Proxemics and Physics-Guided Adaption
- Gaussian Opacity Fields: Efficient Adaptive Surface Reconstruction in Unbounded Scenes
- HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision
- MonoSelfRecon: Purely Self-Supervised Explicit Generalizable 3D Reconstruction of Indoor Scenes from Monocular RGB Views
- Fast Encoder-Based 3D from Casual Videos via Point Track Processing
- 3D Geometry-aware Deformable Gaussian Splatting for Dynamic View Synthesis
- Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind
- SpatialTracker: Tracking Any 2D Pixels in 3D Space
- Per-Gaussian Embedding-Based Deformation for Deformable 3D Gaussian Splatting
- CityGaussian: Real-time High-quality Large-Scale Scene Rendering with Gaussians
- SceneTracker: Long-term Scene Flow Estimation Network
- GlORIE-SLAM: Globally Optimized RGB-only Implicit Encoding Point Cloud SLAM
- InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object Interaction
- Octree-GS: Towards Consistent Real-time Rendering with LOD-Structured 3D Gaussians
- TRAM: Global Trajectory and Motion of 3D Humans from in-the-wild Videos
- Track Everything Everywhere Fast and Robustly
- TC4D: Trajectory-Conditioned Text-to-4D Generation
- EgoLifter: Open-world 3D Segmentation for Egocentric Perception
- GaussianFlow: Splatting Gaussian Dynamics for 4D Content Creation
- SceneScript: Reconstructing Scenes With An Autoregressive Structured Language Model
- SV3D: Novel Multi-view Synthesis and 3D Generation from a Single Image using Latent Video Diffusion
- Recent Advances in 3D Gaussian Splatting
- 3D-VLA: A 3D Vision-Language-Action Generative World Model
- Scaling Up Dynamic Human-Scene Interaction Modeling
- V3D: Video Diffusion Models are Effective 3D Generators
- UFORecon: Generalizable Sparse-View Surface Reconstruction from Arbitrary and UnFavOrable Sets
- DaReNeRF: Direction-aware Representation for Dynamic Scenes
- 3DGStream: On-the-Fly Training of 3D Gaussians for Efficient Streaming of Photo-Realistic Free-Viewpoint Videos
- The Essential Role of Causality in Foundation World Models for Embodied AI
- 4D-Rotor Gaussian Splatting: Towards Efficient Novel View Synthesis for Dynamic Scenes
- A Comprehensive Survey on 3D Content Generation
- Advances in 3D Generation: A Survey
- Spacetime Gaussian Feature Splatting for Real-Time Dynamic View Synthesis
- DUSt3R: Geometric 3D Vision Made Easy
- 3DGS-Avatar: Animatable Avatars via Deformable 3D Gaussian Splatting
- Template Free Reconstruction of Human-object Interaction with Procedural Interaction Generation
- WHAM: Reconstructing World-grounded Humans with Accurate 3D Motion
- Gaussian Splatting SLAM
- I'M HOI: Inertia-aware Monocular Capture of 3D Human-Object Interactions
- PhysHOI: Physics-Based Imitation of Dynamic Human-Object Interaction
- SC-GS: Sparse-Controlled Gaussian Splatting for Editable Dynamic Scenes
- GaussianAvatar: Towards Realistic Human Avatar Modeling from a Single Video via Animatable 3D Gaussians
- GPS-Gaussian: Generalizable Pixel-wise 3D Gaussian Splatting for Real-time Human Novel View Synthesis
- Neural Parametric Gaussians for Monocular Non-Rigid Object Reconstruction
- Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
- Scaffold-GS: Structured 3D Gaussians for View-Adaptive Rendering
- 4D-fy: Text-to-4D Generation Using Hybrid Score Distillation Sampling
- HUGS: Human Gaussian Splats
- Egocentric Whole-Body Motion Capture with FisheyeViT and Diffusion-Based Motion Refinement
- SCALAR-NeRF: SCAlable LARge-scale Neural Radiance Fields for Scene Reconstruction
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- SuGaR: Surface-Aligned Gaussian Splatting for Efficient 3D Mesh Reconstruction and High-Quality Mesh Rendering
- An Embodied Generalist Agent in 3D World
- EmerNeRF: Emergent Spatial-Temporal Scene Decomposition via Self-Supervision
- Real-time Photorealistic Dynamic Scene Representation and Rendering with 4D Gaussian Splatting
- 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering
- Pseudo-Generalized Dynamic View Synthesis from a Video
- Universal Humanoid Motion Representations for Physics-Based Control
- SMPLer-X: Scaling Up Expressive Human Pose and Shape Estimation
- XVO: Generalized Visual Odometry via Cross-Modal Self-Training
- Deformable 3D Gaussians for High-Fidelity Monocular Dynamic Scene Reconstruction
- Physically Plausible Full-Body Hand-Object Interaction Synthesis
- FlowIBR: Leveraging Pre-Training for Efficient Neural Image-Based Rendering of Dynamic Scenes
- GO-SLAM: Global Optimization for Consistent 3D Instant Reconstruction
- Project Aria: A New Tool for Egocentric Multi-Modal AI Research
- Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis
- Guide3D: Create 3D Avatars from Text and Image Guidance
- An Outlook into the Future of Egocentric Vision
- Stereo Visual Odometry with Deep Learning-Based Point and Line Feature Matching using an Attention Graph Neural Network
- AffineGlue: Joint Matching and Robust Estimation
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering
- MAMo: Leveraging Memory and Attention for Monocular Video Depth Estimation
- LightGlue: Local Feature Matching at Light Speed
- C2F2NeUS: Cascade Cost Frustum Fusion for High Fidelity and Generalizable Neural Surface Reconstruction
- Generative Proxemics: A Prior for 3D Social Interaction from Images
- EPIC Fields: Marrying 3D Geometry and Video Understanding
- Tracking Everything Everywhere All at Once
- Neuralangelo: High-Fidelity Neural Surface Reconstruction
- TRACE: 5D Temporal Regression of Avatars with Dynamic Cameras in 3D Environments
- Neural Scene Chronology
- Humans in 4D: Reconstructing and Tracking Humans with Transformers
- Simulation and Retargeting of Complex Multi-Character Interactions
- ReTR: Modeling Rendering Via Transformer for Generalizable Neural Surface Reconstruction
- Reconstructing Animatable Categories from Videos
- Generative AI meets 3D: A Survey on Text-to-3D in AIGC Era
- Perpetual Humanoid Control for Real-time Simulated Avatars
- PMP: Learning to Physically Interact with Environments using Part-wise Motion Priors
- CVRecon: Rethinking 3D Geometric Feature Learning For Neural Reconstruction
- HOSNeRF: Dynamic Human-Object-Scene Neural Radiance Fields from a Single Video
- Instant-3D: Instant Neural Radiance Field Training Towards On-Device AR/VR 3D Reconstruction
- BoDiffusion: Diffusing Sparse Observations for Full-Body Human Motion Synthesis
- VisFusion: Visibility-aware Online 3D Scene Reconstruction from Videos
- Avatars Grow Legs: Generating Smooth Human Motion from Sparse Tracking Inputs with Diffusion Model
- Zip-NeRF: Anti-Aliased Grid-Based Neural Radiance Fields
- Decoupling Dynamic Monocular Videos for Dynamic View Synthesis
- FineRecon: Depth-aware Feed-forward Network for Detailed 3D Reconstruction
- Trace and Pace: Controllable Pedestrian Animation via Guided Trajectory Diffusion
- DreamAvatar: Text-and-Shape Guided 3D Human Avatar Generation via Diffusion Models
- CIRCLE: Capture In Rich Contextual Environments
- 3D Human Pose Estimation via Intuitive Physics
- Flow supervision for Deformable NeRF
- Visibility Aware Human-Object Interaction Tracking from Single RGB Camera
- F2-NeRF: Fast Neural Radiance Field Training with Free Camera Trajectories
- Hi4D: 4D Instance Segmentation of Close Human Interaction
- DyLiN: Making Light Field Networks Dynamic
- NeAT: Learning Neural Implicit Surfaces with Arbitrary Topologies from Multi-view Images
- Decoupling Human and Camera Motion from Videos in the Wild
- Vid2Avatar: 3D Avatar Reconstruction from Videos in the Wild via Self-supervised Scene Decomposition
- NICER-SLAM: Neural Implicit Scene Encoding for RGB SLAM
- K-Planes: Explicit Radiance Fields in Space, Time, and Appearance
- HexPlane: A Fast Representation for Dynamic Scenes
- Dense RGB SLAM with Neural Implicit Maps
- NeRF in the Palm of Your Hand: Corrective Augmentation for Robotics via Novel-View Synthesis
- Robust Dynamic Radiance Fields
- MonoNeRF: Learning a Generalizable Dynamic Radiance Field from Monocular Videos
- Scene-aware Egocentric 3D Human Pose Estimation
- Full-Body Articulated Human-Object Interaction
- Ego-Body Pose Estimation via Ego-Head Pose Estimation
- SPARF: Neural Radiance Fields from Sparse and Noisy Poses
- Tensor4D : Efficient Neural 4D Decomposition for High-fidelity Dynamic Reconstruction and Rendering
- DynIBaR: Neural Dynamic Image-Based Rendering
- Imagen Video: High Definition Video Generation with Diffusion Models
- NeRF: Neural Radiance Field in 3D Vision: A Comprehensive Review (Updated Post-Gaussian Splatting)
- InterCap: Joint Markerless 3D Tracking of Humans and Objects in Interaction
- EPIC-KITCHENS VISOR Benchmark: VIdeo Segmentations and Object Relations
- DytanVO: Joint Refinement of Visual Odometry and Motion Segmentation in Dynamic Environments
- SimpleRecon: 3D Reconstruction Without 3D Convolutions
- MoCapAct: A Multi-Task Dataset for Simulated Humanoid Control
- Deep Patch Visual Odometry
- The One Where They Reconstructed 3D Humans and Environments in TV Shows
- AvatarPoser: Articulated Full-Body Pose Tracking from Sparse Motion Sensing
- Structure PLP-SLAM: Efficient Sparse Mapping and Localization using Point, Line and Plane for Monocular, RGB-D and Stereo Cameras
- Unbiased 4D: Monocular 4D Reconstruction with a Neural Deformation Model
- SparseNeuS: Fast Generalizable Neural Surface Reconstruction from Sparse Views
- FOF: Learning Fourier Occupancy Field for Monocular Real-time Human Reconstruction
- FWD: Real-time Novel View Synthesis with Forward Warping and Depth
- MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface Reconstruction
- Aug-NeRF: Training Stronger Neural Radiance Fields with Triple-Level Physically-Grounded Augmentations
- Capturing and Inferring Dense Full-Body Human-Scene Contact
- DANBO: Disentangled Articulated Neural Body Representations via Graph Neural Networks
- Revealing Occlusions with 4D Neural Fields
- Multi-Frame Self-Supervised Depth with Transformers
- Video Diffusion Models
- Gravitationally Lensed Black Hole Emission Tomography
- CHORE: Contact, Human and Object REconstruction from a single RGB image
- Improving Monocular Visual Odometry Using Learned Depth
- NeuMan: Neural Human Radiance Field from a Single Video
- ϕ-SfT: Shape-from-Template with a Physics-Based Deformation Model
- A Survey of Non-Rigid 3D Registration
- Block-NeRF: Scalable Large Scene Neural View Synthesis
- SelfRecon: Self Reconstruction Your Digital Avatar from Monocular Video
- HumanNeRF: Free-viewpoint Rendering of Moving People from Monocular Video
- BANMo: Building Animatable 3D Neural Models from Many Casual Videos
- Mega-NeRF: Scalable Construction of Large-Scale NeRFs for Virtual Fly-Throughs
- Ref-NeRF: Structured View-Dependent Appearance for Neural Radiance Fields
- GLAMR: Global Occlusion-Aware Human Mesh Recovery with Dynamic Cameras
- Human Performance Capture from Monocular Video in the Wild
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields
- Temporally Consistent Online Depth Estimation in Dynamic Scenes
- Deep Marching Tetrahedra: a Hybrid Representation for High-Resolution 3D Shape Synthesis
- SPEC: Seeing People in the Wild with an Estimated Camera
- Encoder-decoder with Multi-level Attention for 3D Human Shape and Pose Estimation
- Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras
- D3D-HOI: Dynamic 3D Human-Object Interactions from Videos
- Pixel-Perfect Structure-from-Motion with Featuremetric Refinement
- Volume Rendering of Neural Implicit Surfaces
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction
- Dynamic View Synthesis from Dynamic Monocular Video
- Neural Trajectory Fields for Dynamic Novel View Synthesis
- HuMoR: 3D Human Motion Model for Robust Pose Estimation
- Animatable Neural Radiance Fields for Modeling Dynamic Human Bodies
- LASR: Learning Articulated Shape Reconstruction from a Monocular Video
- The Temporal Opportunist: Self-Supervised Multi-Frame Monocular Depth
- H2O: Two Hands Manipulating Objects for First Person Interaction Recognition
- PARE: Part Attention Regressor for 3D Human Body Estimation
- BARF: Bundle-Adjusting Neural Radiance Fields
- Dynamic Surface Function Networks for Clothed Human Bodies
- PhySG: Inverse Rendering with Spherical Gaussians for Physics-based Material Editing and Relighting
- LoFTR: Detector-Free Local Feature Matching with Transformers
- NeuralRecon: Real-Time Coherent 3D Reconstruction from Monocular Video
- FoV-NeRF: Foveated Neural Radiance Fields for Virtual Reality
- PyMAF: 3D Human Pose and Shape Regression with Pyramidal Mesh Alignment Feedback Loop
- In-Place Scene Labelling and Understanding with Implicit Scene Representation
- Generalizing to the Open World: Deep Visual Odometry with Online Adaptation
- KiloNeRF: Speeding up Neural Radiance Fields with Thousands of Tiny MLPs
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields
- Transformer in Transformer
- IBRNet: Learning Multi-View Image-Based Rendering
- A-NeRF: Articulated Neural Radiance Fields for Learning Human Shape, Appearance, and Pose
- Transformers in Vision: A Survey
- Neural Body: Implicit Neural Representations with Structured Latent Codes for Novel View Synthesis of Dynamic Humans
- A Survey on Vision Transformer
- Non-Rigid Neural Radiance Fields: Reconstruction and Novel View Synthesis of a Dynamic Scene From Monocular Video
- End-to-End Human Pose and Mesh Reconstruction with Transformers
- Neural Radiance Flow for 4D View Synthesis and Video Processing
- Robust Consistent Video Depth Estimation
- NeRD: Neural Reflectance Decomposition from Image Collections
- Online Adaptation for Consistent Mesh Reconstruction in the Wild
- pixelNeRF: Neural Radiance Fields from One or Few Images
- PatchmatchNet: Learned Multi-View Patchmatch Stereo
- UniCon: Universal Neural Controller For Physics-based Character Motion
- D-NeRF: Neural Radiance Fields for Dynamic Scenes
- Neural Scene Flow Fields for Space-Time View Synthesis of Dynamic Scenes
- Multi-view Depth Estimation using Epipolar Spatio-Temporal Networks
- Nerfies: Deformable Neural Radiance Fields
- Space-time Neural Irradiance Fields for Free-Viewpoint Video
- GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- NeRF++: Analyzing and Improving Neural Radiance Fields
- Denoising Diffusion Implicit Models
- Monocular, One-stage, Regression of Multiple 3D People
- Human Body Model Fitting by Learned Gradient Descent
- Visibility-aware Multi-view Stereo Network
- Free View Synthesis
- I2L-MeshNet: Image-to-Lixel Prediction Network for Accurate 3D Human Pose and Mesh Estimation from a Single RGB Image
- NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections
- Perceiving 3D Human-Object Spatial Arrangements from a Single Image in the Wild
- Character Controllers Using Motion VAEs
- Long-term Human Motion Prediction with Scene Context
- Differentiable Rendering: A Survey
- Denoising Diffusion Probabilistic Models
- 3D Human Mesh Regression with Dense Correspondence
- Consistent Video Depth Estimation
- Towards Better Generalization: Joint Depth-Pose Learning without PoseNet
- Novel View Synthesis of Dynamic Scenes with Globally Coherent Depths from a Monocular Camera
- Hierarchical Kinematic Human Mesh Recovery
- D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry
- Reformer: The Efficient Transformer
- Don't Forget The Past: Recurrent Depth Estimation from Monocular Video
- Video Depth Estimation by Fusing Flow-to-Depth Proposals
- SynSin: End-to-end View Synthesis from a Single Image
- Cost Volume Pyramid Based Depth Inference for Multi-View Stereo
- Cascade Cost Volume for High-Resolution Multi-View Stereo and Stereo Matching
- VIBE: Video Inference for Human Body Pose and Shape Estimation
- Neural Point Cloud Rendering via Multi-Plane Projection
- SuperGlue: Learning Feature Matching with Graph Neural Networks
- DenseRaC: Joint 3D Pose and Shape Estimation by Dense Render-and-Compare
- Learning to Reconstruct 3D Human Pose and Shape via Model-fitting in the Loop
- DiffTaichi: Differentiable Programming for Physical Simulation
- Temporally Consistent Depth Prediction with Flow-Guided Memory Units
- EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action Recognition
- Resolving 3D Human Pose Ambiguities with 3D Scene Constraints
- Exploiting temporal consistency for real-time video depth estimation
- Unsupervised Scale-consistent Depth and Ego-motion Learning from Monocular Video
- Reinforcement learning
- Self-supervised Learning with Geometric Constraints in Monocular Video: Connecting Flow, Depth, and Camera
- Sim2real transfer learning for 3D human pose estimation: motion to the rescue
- Neural Point-Based Graphics
- Neural-Guided RANSAC: Learning Where to Sample Model Hypotheses
- Convolutional Mesh Regression for Single-Image Human Shape Reconstruction
- Expressive Body Capture: 3D Hands, Face, and Body from a Single Image
- Soft Rasterizer: A Differentiable Renderer for Image-based 3D Reasoning
- DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation
- DeepV2D: Video to Depth with Differentiable Structure from Motion
- Learning 3D Human Dynamics from Video
- LiveCap: Real-time Human Performance Capture from Monocular Video
- Neural Body Fitting: Unifying Deep Learning and Model-Based Human Pose and Shape Estimation
- Mode-adaptive neural networks for quadruped motion control
- Deep Video Portraits
- Learning to Estimate 3D Human Pose and Shape from a Single Color Image
- MVSNet: Depth Inference for Unstructured Multi-view Stereo
- Video Based Reconstruction of 3D People Models
- GeoNet: Unsupervised Learning of Dense Depth, Optical Flow and Camera Pose
- Image Transformer
- SuperPoint: Self-Supervised Interest Point Detection and Description
- End-to-end Recovery of Human Shape and Pose
- Neural 3D Mesh Renderer
- Neural Discrete Representation Learning
- UnDeepVO: Monocular Visual Odometry through Unsupervised Deep Learning
- A Brief Survey of Deep Reinforcement Learning
- Phase-functioned neural networks for character control
- Attention Is All You Need
- A Survey of Structure from Motion
- 3D Menagerie: Modeling the 3D shape and pose of animals
- Keep it SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image
- Direct Sparse Odometry
- Generative Adversarial Imitation Learning
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Screened poisson surface reconstruction
- GauSTAR: Gaussian Surface Tracking and Reconstruction
- Advances in 4D Generation: A Survey
- Dynamic Gaussians Mesh: Consistent Mesh Reconstruction from Dynamic Scenes
- MaGS: Reconstructing and Simulating Dynamic 3D Objects with Mesh-adsorbed Gaussian Splatting
- SkillMimic: Learning Basketball Interaction Skills from Demonstrations
- MoDGS: Dynamic Gaussian Splatting from Casually-captured Monocular Videos with Depth Priors
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Cited by
Related