SPORTS: Simultaneous Panoptic Odometry, Rendering, Tracking and Segmentation for Urban Scenes Understanding
2025/10/14 by Zhiliu Yang, Jinyu Dai, Yang, Zhiliu +5
Computer Science · Environmental Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Remote Sensing and LiDAR Applications #Video Surveillance and Tracking Methods
paper · pdf · doi:10.48550/arxiv.2510.12749
openalex publication_date 2025/10/14 · openalex created_date 2025/10/17 · openalex updated_date 2026/07/28
Abstract
The scene perception, understanding, and simulation are fundamental techniques for embodied-AI agents, while existing solutions are still prone to segmentation deficiency, dynamic objects' interference, sensor data sparsity, and view-limitation problems. This paper proposes a novel framework, named SPORTS, for holistic scene understanding via tightly integrating Video Panoptic Segmentation (VPS), Visual Odometry (VO), and Scene Rendering (SR) tasks into an iterative and unified perspective. Firstly, VPS designs an adaptive attention-based geometric fusion mechanism to align cross-frame features via enrolling the pose, depth, and optical flow modality, which automatically adjust feature maps for different decoding stages. And a post-matching strategy is integrated to improve identities tracking. In VO, panoptic segmentation results from VPS are combined with the optical flow map to improve the confidence estimation of dynamic objects, which enhances the accuracy of the camera pose estimation and completeness of the depth map generation via the learning-based paradigm. Furthermore, the point-based rendering of SR is beneficial from VO, transforming sparse point clouds into neural fields to synthesize high-fidelity RGB views and twin panoptic views. Extensive experiments on three public datasets demonstrate that our attention-based feature fusion outperforms most existing state-of-the-art methods on the odometry, tracking, segmentation, and novel view synthesis tasks.
Citations
- NIS-SLAM: Neural Implicit Semantic RGB-D SLAM for 3D Consistent Scene Understanding
- A Unified Framework for 3D Scene Understanding
- NeRF-XL: Scaling NeRFs with Multiple GPUs
- HUGS: Holistic Urban 3D Scene Understanding via Gaussian Splatting
- DVN-SLAM: Dynamic Visual Neural SLAM Based on Local-Global Encoding
- RoDUS: Robust Decomposition of Static and Dynamic Elements in Urban Scenes
- NiteDR: Nighttime Image De-Raining with Cross-View Sensor Cooperative Learning for Dynamic Driving Scenes
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
- Street Gaussians: Modeling Dynamic Urban Scenes with Gaussian Splatting
- DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Autonomous Driving Scenes
- SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAM
- DGNR: Density-Guided Neural Point Rendering of Large Driving Scenes
- SNI-SLAM: Semantic Neural Implicit SLAM
- Two-Stage Learning of Highly Dynamic Motions with Rigid and Articulated Soft Quadrupeds
- SUDS: Scalable Urban Dynamic Scenes
- S-NeRF: Neural Radiance Fields for Street Views
- NICER-SLAM: Neural Implicit Scene Encoding for RGB SLAM
- DytanVO: Joint Refinement of Visual Odometry and Motion Segmentation in Dynamic Environments
- PVO: Panoptic Visual Odometry
- READ: Large-Scale Neural Scene Rendering for Autonomous Driving
- Video K-Net: A Simple, Strong, and Unified Baseline for Video Segmentation
- Hybrid Tracker with Pixel and Instance for Video Panoptic Segmentation
- Block-NeRF: Scalable Large Scene Neural View Synthesis
- Instant neural graphics primitives with a multiresolution hash encoding
- NICE-SLAM: Neural Implicit Scalable Encoding for SLAM
- Mega-NeRF: Scalable Construction of Large-Scale NeRFs for Virtual Fly-Throughs
- Slot-VPS: Object-centric Representation Learning for Video Panoptic Segmentation
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras
- Rethinking Coarse-to-Fine Approach in Single Image Deblurring
- CTNet: Context-based Tandem Network for Semantic Segmentation
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance\n Fields
- iMAP: Implicit Mapping and Positioning in Real-Time
- ViP-DeepLab: Learning Visual Perception with Depth-aware Video Panoptic Segmentation
- Neural Scene Graphs for Dynamic Scenes
- DOT: Dynamic Object Tracking for Visual SLAM
- ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual–Inertial, and Multimap SLAM
- Quasi-Dense Similarity Learning for Multiple Object Tracking
- Video Panoptic Segmentation
- End-to-End Object Detection with Transformers
- Virtual KITTI 2
- SOLO: Segmenting Objects by Locations
- SimVODIS: Simultaneous Visual Odometry, Object Detection, and Instance Segmentation
- ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks
- ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks
- Neural Point-Based Graphics
- CBAM: Convolutional Block Attention Module
- Playing for Benchmarks
- Squeeze-and-Excitation Networks
- CNN-SLAM: Real-time dense monocular SLAM with learned depth prediction
- Direct Sparse Odometry
- V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation
- Perceptual Losses for Real-Time Style Transfer and Super-Resolution
- Vision meets robotics: The KITTI dataset
- ‘Structure-from-Motion’ photogrammetry: A low-cost, effective tool for geoscience applications
- DDN-SLAM: Real-time Dense Dynamic Neural Implicit SLAM
Related