RangeSAM: On the Potential of Visual Foundation Models for Range-View represented LiDAR segmentation
2025/09/19 by J.H. Kühn, Kühn, Paul Julius, Duc Anh Nguyen +6
Environmental Science · Engineering · Computer Science · #Remote Sensing and LiDAR Applications #Robotics and Sensor-Based Localization #Video Surveillance and Tracking Methods
paper · pdf · doi:10.48550/arxiv.2509.15886
Abstract
Point cloud segmentation is central to autonomous driving and 3D scene understanding. While voxel- and point-based methods dominate recent research due to their compatibility with deep architectures and ability to capture fine-grained geometry, they often incur high computational cost, irregular memory access, and limited real-time efficiency. In contrast, range-view methods, though relatively underexplored - can leverage mature 2D semantic segmentation techniques for fast and accurate predictions. Motivated by the rapid progress in Visual Foundation Models (VFMs) for captioning, zero-shot recognition, and multimodal tasks, we investigate whether SAM2, the current state-of-the-art VFM for segmentation tasks, can serve as a strong backbone for LiDAR point cloud segmentation in the range view. We present , to our knowledge, the first range-view framework that adapts SAM2 to 3D segmentation, coupling efficient 2D feature extraction with standard projection/back-projection to operate on point clouds. To optimize SAM2 for range-view representations, we implement several architectural modifications to the encoder: (1) a novel module that emphasizes horizontal spatial dependencies inherent in LiDAR range images, (2) a customized configuration of tailored to the geometric properties of spherical projections, and (3) an adapted mechanism in the encoder backbone specifically designed to capture the unique spatial patterns and discontinuities present in range-view pseudo-images. Our approach achieves competitive performance on SemanticKITTI while benefiting from the speed, scalability, and deployment simplicity of 2D-centric pipelines. This work highlights the viability of VFMs as general-purpose backbones for 3D perception and opens a path toward unified, foundation-model-driven LiDAR segmentation. Results lets us conclude that range-view segmentation methods using VFMs leads to promising results.
Citations
- AuralSAM2: Enabling SAM2 Hear Through Pyramid Audio-Visual Feature Prompting
- DINO in the Room: Leveraging 2D Foundation Models for 3D Segmentation
- Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
- SAM2-UNet: Segment Anything 2 Makes Strong Encoder for Natural and Medical Image Segmentation
- SAM2-Adapter: Evaluating & Adapting Segment Anything 2 in Downstream Tasks: Camouflage, Shadow, Medical Image Segmentation, and More
- A Closer Look at Deep Learning Methods on Tabular Datasets
- EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
- Point Transformer V3: Simpler, Faster, Stronger
- Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks
- Towards Large-scale 3D Representation Learning with Multi-dataset Point Prompt Training
- Semantic-SAM: Segment and Recognize Anything at Any Granularity
- Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles
- DINOv2: Learning Robust Visual Features without Supervision
- Segment Everything Everywhere All at Once
- Micrograph segmentations for DDEVD
- Segment Anything
- EVA-CLIP: Improved Training Techniques for CLIP at Scale
- EVA-02: A visual representation for neon genesis
- Annotated Point Clouds and Images from a Spatial-Semantic Perception Pipeline: Robotized Deconstruction at the Reference Construction Site, Aachen
- Rethinking Range View Representation for LiDAR Segmentation
- CLIP-FO3D: Learning Free Open-world 3D Scene Representations from 2D Dense CLIP
- Jaccard Metric Losses: Optimizing the Jaccard Index with Soft Labels
- RangeViT: Towards Vision Transformers for 3D Semantic Segmentation in Autonomous Driving
- CLIP2Scene: Towards Label-efficient 3D Scene Understanding by CLIP
- OpenScene: 3D Scene Understanding with Open Vocabularies
- EVA: Exploring the Limits of Masked Visual Representation Learning at Scale
- Point Transformer V2: Grouped Vector Attention and Partition-based Pooling
- CENet: Toward Concise and Efficient LiDAR Semantic Segmentation for Autonomous Driving
- MaskRange: A Mask-classification Model for Range-view based LiDAR Segmentation
- Point-to-Voxel Knowledge Distillation for LiDAR Semantic Segmentation
- Florence: A New Foundation Model for Computer Vision
- On the Opportunities and Risks of Foundation Models
- Emerging Properties in Self-Supervised Vision Transformers
- Emerging Properties in Self-Supervised Vision Transformers
- Lite-HDSeg: LiDAR Semantic Segmentation Using Lite Harmonic Dense Convolutions
- 3D Object Detection with Pointformer
- Multi Projection Fusion for Real-time Semantic Segmentation of 3D LiDAR Point Clouds
- Cylinder3D: An Effective 3D Framework for Driving-scene LiDAR Semantic Segmentation
- Searching Efficient 3D Architectures with Sparse Point-Voxel Convolution
- KPRNet: Improving projection-based LiDAR semantic segmentation
- PolarNet: An Improved Grid Representation for Online LiDAR Point Clouds Semantic Segmentation
- 3D-MPA: Multi Proposal Aggregation for 3D Semantic Instance Segmentation
- SalsaNext: Fast, Uncertainty-aware Semantic Segmentation of LiDAR Point Clouds for Autonomous Driving
- PointASNL: Robust Point Clouds Processing using Nonlocal Neural Networks with Adaptive Sampling
- SemanticPOSS: A Point Cloud Dataset with Large Quantity of Dynamic Instances
- Scalability in Perception for Autonomous Driving: Waymo Open Dataset
- Scalability in Perception for Autonomous Driving: Waymo Open Dataset
- RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds
- RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds
- LATTE: Accelerating LiDAR Point Cloud Annotation via Sensor Fusion, One-Click Annotation, and Tracking
- KPConv: Flexible and Deformable Convolution for Point Clouds
- JSIS3D: Joint Semantic-Instance Segmentation of 3D Point Clouds with Multi-Task Pointwise Networks and Multi-Value Conditional Random Fields
- nuScenes: A multimodal dataset for autonomous driving
- Post Processing of image segmentation using Conditional Random Fields
- PointConv: Deep Convolutional Networks on 3D Point Clouds
- A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation
- SqueezeSegV2: Improved Model Structure and Unsupervised Domain Adaptation for Road-Object Segmentation from a LiDAR Point Cloud
- Dynamic Graph CNN for Learning on Point Clouds
- Dynamic Graph CNN for Learning on Point Clouds
- Large-scale Point Cloud Semantic Segmentation with Superpoint Graphs
- Frustum PointNets for 3D Object Detection from RGB-D Data
- Receptive Field Block Net for Accurate and Fast Object Detection
- SqueezeSeg: Convolutional Neural Nets with Recurrent CRF for Real-Time Road-Object Segmentation from 3D LiDAR Point Cloud
- Attention Is All You Need
- PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space
- PointNet: Deep Learning on Point Sets for 3D Classification and\n Segmentation
- FractalNet: Ultra-Deep Neural Networks without Residuals
- The Cityscapes Dataset for Semantic Urban Scene Understanding
- U-Net: Convolutional Networks for Biomedical Image Segmentation
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Related