ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
2017/02/14 by Angela Dai, Dai, Angela, Manolis Savva +8 · 419 citations
Engineering · Computer Science · Earth and Planetary Sciences · #Robotics and Sensor-Based Localization #Advanced Vision and Imaging #3D Surveying and Cultural Heritage
paper · pdf · doi:10.48550/arxiv.1702.04405
Abstract
A key requirement for leveraging supervised deep learning methods is the availability of large, labeled datasets. Unfortunately, in the context of RGB-D scene understanding, very little data is available -- current datasets cover a small range of scene views and have limited semantic annotations. To address this issue, we introduce ScanNet, an RGB-D video dataset containing 2.5M views in 1513 scenes annotated with 3D camera poses, surface reconstructions, and semantic segmentations. To collect this data, we designed an easy-to-use and scalable RGB-D capture system that includes automated surface reconstruction and crowdsourced semantic annotation. We show that using this data helps achieve state-of-the-art performance on several 3D scene understanding tasks, including 3D object classification, semantic voxel labeling, and CAD model retrieval. The dataset is freely available at http://www.scan-net.org.
Citations
Cited by
- GVSynergy-Det: Synergistic Gaussian-Voxel Representations for Multi-View 3D Object Detection
- 3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds
- Split4D: Decomposed 4D Scene Reconstruction Without Video Segmentation
- ShapeR: Robust Conditional 3D Shape Generation from Casual Captures
- Visual Autoregressive Modelling for Monocular Depth Estimation
- MEGA-PCC: A Mamba-based Efficient Approach for Joint Geometry and Attribute Point Cloud Compression
- VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement
- Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation
- RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
- GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding
- FARM: Find Anything using Relational Spatial Memory
- PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud Learning
- Geometric Context Transformer for Streaming 3D Reconstruction
- ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training
- Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow
- Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer
- Analyzing the Mechanism of Attention Collapse in VGGT from a Dynamics Perspective
- Quantile Rendering: Efficiently Embedding High-dimensional Feature on 3D Gaussian Splatting
- Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
- MVInverse: Feed-forward Multi-view Inverse Rendering in Seconds
- PUFM++: Point Cloud Upsampling via Enhanced Flow Matching
- PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding
- From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
- G3Splat: Geometrically Consistent Generalizable Gaussian Splatting
- FLEG: Feed-Forward Language Embedded Gaussian Splatting from Any Views via Compact Semantic Representation
- 4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation
- Scaling Spatial Reasoning in MLLMs through Programmatic Data Synthesis
- SegGraph: Leveraging Graphs of SAM Segments for Few-Shot 3D Part Segmentation
- Auto-Vocabulary 3D Object Detection
- From Theory to Throughput: CUDA-Optimized APML for Large-Batch 3D Learning
- Multi-View Foundation Models
- MoonSeg3R: Monocular Online Zero-Shot Segment Anything in 3D with Reconstructive Foundation Priors
- SemanticBridge - A Dataset for 3D Semantic Segmentation of Bridges and Domain Gap Analysis
- Off The Grid: Detection of Primitives for Feed-Forward 3D Gaussian Splatting
- Unified Semantic Transformer for 3D Scene Understanding
- LitePT: Lighter Yet Stronger Point Transformer
- Recurrent Video Masked Autoencoders
- LASER: Layer-wise Scale Alignment for Training-Free Streaming 4D Reconstruction
- Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos
- D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation
- Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video
- DOS: Distilling Observable Softmaps of Zipfian Prototypes for Self-Supervised Point Representation
- SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Model
- MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence
- Video Depth Propagation
- XDen-1K: A Density Field Dataset of Real-World Objects
- Long-LRM++: Preserving Fine Details in Feed-Forward Wide-Coverage Reconstruction
- D2GSLAM: 4D Dynamic Gaussian Splatting SLAM
- Geometry-to-Image Synthesis-Driven Generative Point Cloud Registration
- FUSER: Feed-Forward MUltiview 3D Registration Transformer and SE(3)N Diffusion Refinement
- ASSIST-3D: Adapted Scene Synthesis for Class-Agnostic 3D Instance Segmentation
- GimbalDiffusion: Gravity-Aware Camera Control for Video Generation
- SIP: Site in Pieces- A Dataset of Disaggregated Construction-Phase 3D Scans for Semantic Segmentation and Scene Understanding
- Efficiently Reconstructing Dynamic Scenes One D4RT at a Time
- OpenMonoGS-SLAM: Monocular Gaussian Splatting SLAM with Open-set Semantics
- Query-aware Hub Prototype Learning for Few-Shot 3D Point Cloud Semantic Segmentation
- CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning
- D3-Predictor: Noise-Free Deterministic Diffusion for Dense Prediction
- Online Segment Any 3D Thing as Instance Tracking
- Hierarchical Image-Guided 3D Point Cloud Segmentation in Industrial Scenes via Multi-View Bayesian Fusion
- Joint 3D Geometry Reconstruction and Motion Generation for 4D Synthesis from a Single Image
- LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging
- Parabolic Position Encoding: Vision-Centric, Principled, Extrapolatable, General
- Towards Cross-View Point Correspondence in Vision-Language Models
- When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering
- StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
- Unique Lives, Shared World: Learning from Single-Life Videos
- C3G: Learning Compact 3D Representations with 2K Gaussians
- ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
- OpenTrack3D: Towards Accurate and Generalizable Open-Vocabulary 3D Instance Segmentation
- What Is The Best 3D Scene Representation for Robotics? From Geometric to Foundation Models
- Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
- GraphFusion3D: Dynamic Graph Attention Convolution with Adaptive Cross-Modal Transformer for 3D Object Detection
- AVGGT: Rethinking Global Attention for Accelerating VGGT
- Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding
- SpatialReasoner: Active Perception for Large-Scale 3D Scene Understanding
- Register Any Point: Scaling 3D Point Cloud Registration by Flow Matching
- FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention
- InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
- FOD-S2R: A FOD Dataset for Sim2Real Transfer Learning based Object Detection
- S2-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
- Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model
- LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency
- DenseScan: Advancing 3D Scene Understanding with 2D Dense Annotation
- Robust 3DGS-based SLAM via Adaptive Kernel Smoothing
- SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models
- Taming the Light: Illumination-Invariant Semantic 3DGS-SLAM
- ViGG: Robust RGB-D Point Cloud Registration using Visual-Geometric Mutual Guidance
- MARVO: Marine-Adaptive Radiance-aware Visual Odometry
- Fin3R: Fine-tuning Feed-forward 3D Reconstruction Models via Monocular Knowledge Distillation
- Seeing without Pixels: Perception from Camera Trajectories
- Surface Normal Estimation of Tilted Images via Spatial Rectifier
- Resolution Where It Counts: Hash-based GPU-Accelerated 3D Reconstruction via Variance-Adaptive Voxel Grids
- HTTM: Head-wise Temporal Token Merging for Faster VGGT
- Unlocking Zero-shot Potential of Semi-dense Image Matching via Gaussian Splatting
- Scenes as Tokens: Multi-Scale Normal Distributions Transform Tokenizer for General 3D Vision-Language Understanding
- G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
- MODEST: Multi-Optics Depth-of-Field Stereo Dataset
- Accelerating Sparse Convolutions in Voxel-Based Point Cloud Networks
- Vision-Language Memory for Spatial Reasoning
- Wanderland: Geometrically Grounded Simulation for Open-World Embodied AI
- AudioScene: Integrating Object-Event Audio into 3D Scenes
- DAPointMamba: Domain Adaptive Point Mamba for Point Cloud Completion
- Zoo3D: Zero-Shot 3D Object Detection at Scene Level
- Foundry: Distilling 3D Foundation Models for the Edge
- AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with Backend
- RADSeg: Unleashing Parameter and Compute Efficient Zero-Shot Open-Vocabulary Segmentation Using Agglomerative Models
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- FilmSceneDesigner: Chaining Set Design for Procedural Film Scene Generation
- DetAny4D: Detect Anything 4D Temporally in a Streaming RGB Video
- Edit2Perceive: Image Editing Diffusion Models Are Strong Dense Perceivers
- Disc3D: Automatic Curation of High-Quality 3D Dialog Data via Discriminative Object Referring
- 4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation
- EventBench: Towards Comprehensive Benchmarking of Event-based MLLMs
- SING3R-SLAM: Submap-based Indoor Monocular Gaussian SLAM with 3D Reconstruction Priors
- Late-decoupled 3D Hierarchical Semantic Segmentation with Semantic Prototype Discrimination based Bi-branch Supervision
- POMA-3D: The Point Map Way to 3D Scene Understanding
- YOWO: You Only Walk Once to Jointly Map An Indoor Scene and Register Ceiling-mounted Cameras
- Real-Time 3D Object Detection with Inference-Aligned Learning
- LLaVA3: Representing 3D Scenes like a Cubist Painter to Boost 3D Scene Understanding of VLMs
- LEGO-SLAM: Language-Embedded Gaussian Optimization SLAM
- RoMa v2: Harder Better Faster Denser Feature Matching
- SweeperBot: Making 3D Browsing Accessible through View Analysis and Visual Question Answering
- Video Spatial Reasoning with Object-Centric 3D Rollout
- DiffPixelFormer: Differential Pixel-Aware Transformer for RGB-D Indoor Scene Segmentation
- Scan2CAD: Learning CAD Model Alignment in RGB-D Scans
- CLIDD: Cross-Layer Independent Deformable Description for Efficient and Discriminative Local Feature Representation
- DoReMi: A Domain-Representation Mixture Framework for Generalizable 3D Understanding
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods
- MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation
- DBGroup: Dual-Branch Point Grouping for Weakly Supervised 3D Semantic Instance Segmentation
- IPCD: Intrinsic Point-Cloud Decomposition
- STORM: Segment, Track, and Object Re-Localization from a Single Image
- EPSegFZ: Efficient Point Cloud Semantic Segmentation for Few- and Zero-Shot Scenarios with Language Guidance
- Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy
- TrueCity: Real and Simulated Urban Data for Cross-Domain 3D Scene Understanding
- Structured3D: A Large Photo-realistic Dataset for Structured 3D Modeling
- PointASNL: Robust Point Clouds Processing using Nonlocal Neural Networks with Adaptive Sampling
- MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
- U(PM)2:Unsupervised polygon matching with pre-trained models for challenging stereo images
- Point Cloud Segmentation of Integrated Circuits Package Substrates Surface Defects Using Causal Inference: Dataset Construction and Methodology
- How Many Tokens Do 3D Point Cloud Transformer Architectures Really Need?
- iFlyBot-VLM Technical Report
- Visual Spatial Tuning
- Cambrian-S: Towards Spatial Supersensing in Video
- Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts
- Room Envelopes: A Synthetic Dataset for Indoor Layout Reconstruction from Images
- OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic Camera
- Dynamic Reflections: Probing Video Representations with Text Alignment
- 3EED: Ground Everything Everywhere in 3D
- Class-agnostic 3D Segmentation by Granularity-Consistent Automatic 2D Mask Tracking
- Benchmarking individual tree segmentation using multispectral airborne laser scanning data: the FGI-EMIT dataset
- Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models
- Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
- 3D Guided Weakly Supervised Semantic Segmentation
- Towards Part-Based Understanding of RGB-D Scans
- PSTNet: Point Spatio-Temporal Convolution on Point Cloud Sequences
- Glob3R: Global Structure-from-Motion with 3D Foundation Models
- Consistent Video Depth Estimation
- Shape Inpainting using 3D Generative Adversarial Network and Recurrent Convolutional Networks
- CG-World: A Large-Scale World-State Dataset and Protocol for World Models
- Deep point embedding for urban classification using ALS point clouds: A new perspective from local to global
- Point Attention Network for Semantic Segmentation of 3D Point Clouds
- Learning Camera Localization via Dense Scene Matching
- H3DNet: 3D Object Detection Using Hybrid Geometric Primitives
- Deep Robust Single Image Depth Estimation Neural Network Using Scene Understanding
- Multi-view Depth Estimation using Epipolar Spatio-Temporal Networks
- Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning
- AerialMetric: Benchmarking and Adapting UAV Monocular Metric Depth Estimation in the Real World
- PixARMesh: Autoregressive Mesh-Native Single-View Scene Reconstruction
- Are We Hungry for 3D LiDAR Data for Semantic Segmentation? A Survey and Experimental Study
- STD: Sparse-to-Dense 3D Object Detector for Point Cloud
- VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
- The StreetLearn Environment and Dataset
- Robust Consistent Video Depth Estimation
- Loop Closure Detection with RGB-D Feature Pyramid Siamese Networks
- Towards Semantic Segmentation of Urban-Scale 3D Point Clouds: A Dataset, Benchmarks and Challenges
- RaCo: Ranking and Covariance for Practical Learned Keypoints
- 3DIoUMatch: Leveraging IoU Prediction for Semi-Supervised 3D Object Detection
- XRefine: Attention-Guided Keypoint Match Refinement
- EA3D: Online Open-World 3D Object Extraction from Streaming Videos
- AtlasGS: Atlanta-world Guided Surface Reconstruction with Implicit Structured Gaussians
- 3D Annotation Of Arbitrary Objects In The Wild
- LightGlueStick: a Fast and Robust Glue for Joint Point-Line Matching
- Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
- More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models
- Explicit Memory through Online 3D Gaussian Splatting Improves Class-Agnostic Video Segmentation
- PointGroup: Dual-Set Point Grouping for 3D Instance Segmentation
- Scaling Up Occupancy-centric Driving Scene Generation: Dataset and Method
- Gen-LangSplat: Generalized Language Gaussian Splatting with Pre-Trained Feature Compression
- IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction
- MOGRAS: Human Motion with Grasping in 3D Scenes
- DAP-MAE: Domain-Adaptive Point Cloud Masked Autoencoder for Effective Cross-Domain Learning
- PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments
- AnyPcc: Compressing Any Point Cloud with a Single Universal Model
- COS3D: Collaborative Open-Vocabulary 3D Segmentation
- AccSS3D: Accelerator for Spatially Sparse 3D DNNs
- PoseCrafter: Extreme Pose Estimation with Hybrid Video Synthesis
- AegisRF: Adversarial Perturbations Guided with Sensitivity for Protecting Intellectual Property of Neural Radiance Fields
- OpenInsGaussian: Open-vocabulary Instance Gaussian Segmentation with Context-aware Cross-view Fusion
- PLANA3R: Zero-shot Metric Planar 3D Reconstruction via Feed-Forward Planar Splatting
- HouseTour: A Virtual Real Estate A(I)gent
- 3D Weakly Supervised Semantic Segmentation via Class-Aware and Geometry-Guided Pseudo-Label Refinement
- Towards 3D Objectness Learning in an Open World
- GSPlane: Concise and Accurate Planar Reconstruction via Structured Representation
- GuideFlow3D: Optimization-Guided Rectified Flow For Appearance Transfer
- SaLon3R: Structure-aware Long-term Generalizable 3D Reconstruction from Unposed Images
- Terra: Explorable Native 3D World Model with Point Latents
- ChangingGrounding: 3D Visual Grounding in Changing Scenes
- C4D: 4D Made from 3D through Dual Correspondences
- Leveraging Cycle-Consistent Anchor Points for Self-Supervised RGB-D Registration
- PU-Transformer: Point Cloud Upsampling Transformer
- VIST3A: Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator
- DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
- Reasoning in Space via Grounding in the World
- BEEP3D: Box-Supervised End-to-End Pseudo-Mask Generation for 3D Instance Segmentation
- Scene Coordinate Reconstruction Priors
- IL3D: A Large-Scale Indoor Layout Dataset for LLM-Driven 3D Scene Generation
- ACE-G: Improving Generalization of Scene Coordinate Regression Through Query Pre-Training
- SNAP: Towards Segmenting Anything in Any Point Cloud
- Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
- Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph
- Visual Odometry with Transformers
- GeoPurify: A Data-Efficient Geometric Distillation Framework for Open-Vocabulary 3D Segmentation
- WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting
- EC3R-SLAM: Efficient and Consistent Monocular Dense SLAM with Feed-Forward 3D Reconstruction
- From Programs to Poses: Factored Real-World Scene Generation via Learned Program Libraries
- SpaceVista: All-Scale Visual Spatial Reasoning from mm to km
- Visibility-Aware Densification for 3D Gaussian Splatting in Dynamic Urban Scenes
- SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
- AlignGS: Aligning Geometry and Semantics for Robust Indoor Reconstruction from Sparse Views
- Robust 2D/3D Vehicle Parsing in CVIS
- ARTDECO: Towards Efficient and High-Fidelity On-the-Fly 3D Reconstruction with Structured Scene Representation
- Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
- MoRe: Monocular Geometry Refinement via Graph Optimization for Cross-View Consistency
- UniFField: A Generalizable Unified Neural Feature Field for Visual, Semantic, and Spatial Uncertainties in Any Scene
- Geometry-Aware Cross Modal Alignment for Light Field-LiDAR Semantic Segmentation
- RTGS: Real-Time 3D Gaussian Splatting SLAM via Multi-Level Redundancy Reduction
- Mangrove3D: Terrestrial Laser Scanning Dataset for Coastal Mangrove Forests
- Diffusion2: Turning 3D Environments into Radio Frequency Heatmaps
- Improved probabilistic regression using diffusion models
- 3D-SIS: 3D Semantic Instance Segmentation of RGB-D Scans
- Learning the Depths of Moving People by Watching Frozen People
- GS-Share: Enabling High-fidelity Map Sharing with Incremental Gaussian Splatting
- Spatial-ViLT: Enhancing Visual Spatial Reasoning through Multi-Task Learning
- KeySG: Hierarchical Keyframe-Based 3D Scene Graphs
- Semantic Visual Simultaneous Localization and Mapping: A Survey on State of the Art, Challenges, and Future Directions
- PointConv: Deep Convolutional Networks on 3D Point Clouds
- PhraseStereo: The First Open-Vocabulary Stereo Image Segmentation Dataset
- DEPTHOR++: Robust Depth Enhancement from a Real-World Lightweight dToF and RGB Guidance
- Text-to-Scene with Large Reasoning Models
- Image-Plane Geometric Decoding for View-Invariant Indoor Scene Reconstruction
- PinPoint3D: Fine-Grained 3D Part Segmentation from a Few Clicks
- TTT3R: 3D Reconstruction as Test-Time Training
- LLM-RG: Referential Grounding in Outdoor Scenarios using Large Language Models
- BRIDGE -- Building Reinforcement-Learning Depth-to-Image Data Generation Engine for Monocular Depth Estimation
- Social 3D Scene Graphs: Modeling Human Actions and Relations for Interactive Service Robots
- NeuralPVS: Learned Estimation of Potentially Visible Sets
- CORE-3D: Context-aware Open-vocabulary Retrieval by Embeddings in 3D
- Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning Synergy
- OMeGa: Joint Optimization of Explicit Meshes and Gaussian Splats for Robust Scene-Level Surface Reconstruction
- LiDAR-based Panoptic Segmentation via Dynamic Shifting Network
- What can I do here? Leveraging Deep 3D saliency and geometry for fast and scalable multiple affordance detection
- Multiview Based 3D Scene Understanding On Partial Point Sets
- Euclid's Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks
- GRS-SLAM3R: Real-Time Dense SLAM with Gated Recurrent State
- InteractMove: Text-Controlled Human-Object Interaction Generation in 3D Scenes with Movable Objects
- OVSeg3R: Learn Open-vocabulary Instance Segmentation from 2D via 3D Reconstruction
- M3DLayout: A Multi-Source Dataset of 3D Indoor Layouts and Structured Descriptions for 3D Generation
- GeLoc3r: Enhancing Relative Camera Pose Regression with Geometric Consistency Regularization
- Polysemous Language Gaussian Splatting via Matching-based Mask Lifting
- Joint graph entropy knowledge distillation for point cloud classification and robustness against corruptions
- FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D Prediction
- BIDCD -- Bosch Industrial Depth Completion Dataset
- Shape Completion using 3D-Encoder-Predictor CNNs and Shape Synthesis
- Reflect3r: Single-View 3D Stereo Reconstruction Aided by Mirror Reflections
- A Mobile Manipulation System for One-Shot Teaching of Complex Tasks in Homes
- Convolutional Occupancy Networks
- TUN3D: Towards Real-World Scene Understanding from Unposed Images
- Articulated Object Reconstruction from Rest-State Observation
- 3D Objectness Estimation via Bottom-up Regret Grouping
- ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA
- 3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single Image
- CAM-Convs: Camera-Aware Multi-Scale Convolutions for Single-View Depth
- RevealNet: Seeing Behind Objects in RGB-D Scans
- AnyDepth: Depth Estimation Made Easy
- LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
- DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimation
- Relational Scene Graphs for Object Grounding of Natural Language Commands
- Learning Equivariant Representations
- Point Cloud Instance Segmentation with Semi-supervised Bounding-Box Mining
- Self-Supervised Learning for Domain Adaptation on Point-Clouds
- ParaNet: Deep Regular Representation for 3D Point Clouds
- SGAligner++: Cross-Modal Language-Aided 3D Scene Graph Alignment
- RS3DBench: A Comprehensive Benchmark for 3D Spatial Perception in Remote Sensing
- VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction
- SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
- 3DContextNet: K-d Tree Guided Hierarchical Learning of Point Clouds Using Local and Global Contextual Cues
- \mathttM3VIR: A Large-Scale Multi-Modality Multi-View Synthesized Benchmark Dataset for Image Restoration and Content Creation
- ConfidentSplat: Confidence-Weighted Depth Fusion for Accurate 3D Gaussian Splatting SLAM
- SLAM-Former: Putting SLAM into One Transformer
- NUMINA: A Natural Understanding Benchmark for Multi-dimensional Intelligence and Numerical Reasoning Abilities
- Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
- SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D Features
- See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
- MeshSplat: Generalizable Sparse-View Surface Reconstruction via Gaussian Splatting
- 3D Gaussian Flats: Hybrid 2D/3D Photometric Scene Reconstruction
- Sparse Multiview Open-Vocabulary 3D Detection
- Investigating Domain Gaps for Indoor 3D Object Detection
- SCENEFORGE: Enhancing 3D-text alignment with Structured Scene Compositions
- CAGE: Continuity-Aware edGE Network Unlocks Robust Floorplan Reconstruction
- PlaneSegNet: Fast and Robust Plane Estimation Using a Single-stage Instance Segmentation CNN
- SPATIALGEN: Layout-guided 3D Indoor Scene Generation
- Efficient 3D Perception on Embedded Systems via Interpolation-Free Tri-Plane Lifting and Volume Fusion
- Survey on semantic segmentation using deep learning techniques
- Deep Learning for 3D Point Cloud Understanding: A Survey
- BIM Informed Visual SLAM for Construction Monitoring
- White Aggregation and Restoration for Few-shot 3D Point Cloud Semantic Segmentation
- ScanComplete: Large-Scale Scene Completion and Semantic Segmentation for 3D Scans
- TextureNet: Consistent Local Parametrizations for Learning from High-Resolution Signals on Meshes
- Few to Big: Prototype Expansion Network via Diffusion Learner for Point Cloud Few-shot Semantic Segmentation
- UDON: Uncertainty-weighted Distributed Optimization for Multi-Robot Neural Implicit Mapping under Extreme Communication Constraints
- SimVODIS: Simultaneous Visual Odometry, Object Detection, and Instance Segmentation
- 3D Aware Region Prompted Vision Language Model
- PointMixer: MLP-Mixer for Point Cloud Understanding
- OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
- 3DAeroRelief: The first 3D Benchmark UAV Dataset for Post-Disaster Assessment
- M3DMap: Object-aware Multimodal 3D Mapping for Dynamic Environments
- InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts
- Few-shot 3D Point Cloud Semantic Segmentation
- LayerLock: Non-collapsing Representation Learning with Progressive Freezing
- Towards Understanding Visual Grounding in Visual Language Models
- Generalized Zero-Shot Learning for Point Cloud Segmentation with Evidence-Based Dynamic Calibration
- PointContrast: Unsupervised Pre-training for 3D Point Cloud Understanding
- Point Linguist Model: Segment Any Object via Bridged Large 3D-Language Model
- EDFFDNet: Towards Accurate and Efficient Unsupervised Multi-Grid Image Registration
- Multi-Path Region Mining For Weakly Supervised 3D Semantic Segmentation on Point Clouds
- APML: Adaptive Probabilistic Matching Loss for Robust 3D Point Cloud Reconstruction
- P3-SAM: Native 3D Part Segmentation
- Towards scalable organ level 3D plant segmentation: Bridging the data algorithm computing gap
- Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers
- Guided Point Contrastive Learning for Semi-supervised Point Cloud Semantic Segmentation
- Shape from Shading through Shape Evolution
- PanopticFusion: Online Volumetric Semantic Mapping at the Level of Stuff and Things
- JRN-Geo: A Joint Perception Network based on RGB and Normal images for Cross-view Geo-localization
- Visibility-Aware Language Aggregation for Open-Vocabulary Segmentation in 3D Gaussian Splatting
- WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool
- SGS-3D: High-Fidelity 3D Instance Segmentation via Reliable Semantic Mask Splitting and Growing
- Towards Open World Detection: A Survey
- From Editor to Dense Geometry Estimator
- SGPN: Similarity Group Proposal Network for 3D Point Cloud Instance Segmentation
- GSPN: Generative Shape Proposal Network for 3D Instance Segmentation in Point Cloud
- OccuSeg: Occupancy-aware 3D Instance Segmentation
- DyCo3D: Robust Instance Segmentation of 3D Point Clouds through Dynamic Convolution
- HAMSt3R: Human-Aware Multi-view Stereo 3D Reconstruction
- Reg3D: Reconstructive Geometry Instruction Tuning for 3D Scene Understanding
- Sparse Single Sweep LiDAR Point Cloud Segmentation via Learning Contextual Shape Priors from Scene Completion
- Compositional Prototype Network with Multi-view Comparision for Few-Shot Point Cloud Semantic Segmentation
- ViSTA-SLAM: Visual SLAM with Symmetric Two-view Association
- Non-local RoIs for Instance Segmentation
- OpenMulti: Open-Vocabulary Instance-Level Multi-Agent Distributed Implicit Mapping
- Robix: A Unified Model for Robot Interaction, Reasoning and Planning
- FGO-SLAM: Enhancing Gaussian SLAM with Globally Consistent Opacity Radiance Field
- RfD-Net: Point Scene Understanding by Semantic Instance Reconstruction
- Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment
- Complete Gaussian Splats from a Single Image with Denoising Diffusion Models
- UnPose: Uncertainty-Guided Diffusion Priors for Zero-Shot Pose Estimation
- ActLoc: Learning to Localize on the Move via Active Viewpoint Selection
- OpenM3D: Open Vocabulary Multi-view Indoor 3D Object Detection without Human Annotations
- Integrating SAM Supervision for 3D Weakly Supervised Point Cloud Segmentation
- DeepThink3D: Enhancing Large Language Models with Programmatic Reasoning in Complex 3D Situated Reasoning Tasks
- MASC: Multi-scale Affinity with Sparse Convolution for 3D Instance Segmentation
- GALA: Guided Attention with Language Alignment for Open Vocabulary Gaussian Splatting
- Online 3D Gaussian Splatting Modeling with Novel View Selection
- ROVR-Open-Dataset: A Large-Scale Depth Dataset for Autonomous Driving
- Training Deep Neural Networks to Detect Repeatable 2D Features Using Large Amounts of 3D World Capture Data
- Instance Segmentation in 3D Scenes using Semantic Superpoint Tree Networks
- G-CUT3R: Guided 3D Reconstruction with Camera and Depth Prior Integration
- Advancing 3D Scene Understanding with MV-ScanQA Multi-View Reasoning Evaluation and TripAlign Pre-training Dataset
- STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer
- A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation
- RayletDF: Raylet Distance Fields for Generalizable 3D Surface Reconstruction from Point Clouds or Gaussians
- Contextual Scene Augmentation and Synthesis via GSACNet
- CitySeg: A 3D Open Vocabulary Semantic Segmentation Foundation Model in City-scale Scenarios
- Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
- SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
- Dream4D: Lifting Camera-Controlled I2V towards Spatiotemporally Consistent 4D Generation
- Understanding Dynamic Scenes in Ego Centric 4D Point Clouds
- Indoor Panorama Planar 3D Reconstruction via Divide and Conquer
- Normal Assisted Stereo Depth Estimation
- ETA: Energy-based Test-time Adaptation for Depth Completion
- Cross-View Localization via Redundant Sliced Observations and A-Contrario Validation
- B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding
- EndoMatcher: Generalizable Endoscopic Image Matcher via Multi-Domain Pre-training for Robot-Assisted Surgery
- A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
- Propagating Sparse Depth via Depth Foundation Model for Out-of-Distribution Depth Completion
- Open-world Point Cloud Semantic Segmentation: A Human-in-the-loop Framework
- AR as an Evaluation Playground: Bridging Metrics and Visual Perception of Computer Vision Models
- Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Images
- Qwen-Image Technical Report
- SGAD: Semantic and Geometric-aware Descriptor for Local Feature Matching
- GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
- OpenGS-Fusion: Open-Vocabulary Dense Mapping with Hybrid 3D Gaussian Splatting for Refined Object-Level Understanding
- Multimodal Referring Segmentation: A Survey
- Cross-Dataset Semantic Segmentation Performance Analysis: Unifying NIST Point Cloud City Datasets for 3D Deep Learning
- 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection
- Stable-Sim2Real: Exploring Simulation of Real-Captured 3D Data with Two-Stage Depth Diffusion
- FastPoint: Accelerating 3D Point Cloud Model Inference via Sample Point Distance Prediction
- VMatcher: State-Space Semi-Dense Local Feature Matching
- Details Matter for Indoor Open-vocabulary 3D Instance Segmentation
- SPARE3D: A Dataset for SPAtial REasoning on Three-View Line Drawings
- Modality-Aware Feature Matching: A Comprehensive Review of Single- and Cross-Modality Techniques
- DI-Fusion: Online Implicit 3D Reconstruction with Deep Priors
- Graph-Guided Dual-Level Augmentation for 3D Scene Segmentation
- Multi-view PointNet for 3D Scene Understanding
- UAVScenes: A Multi-Modal Dataset for UAVs
- Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos
- FPConv: Learning Local Flattening for Point Convolution
- SK-Net: Deep Learning on Point Cloud via End-to-end Discovery of Spatial Keypoints
Related