What Demands Attention in Urban Street Scenes? From Scene Understanding towards Road Safety: A Survey of Vision-driven Datasets and Studies
2025/07/09 by Yaoqi Huang, Julie Stephany Berrío, Huang, Yaoqi +5
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Video Surveillance and Tracking Methods
paper · pdf · doi:10.48550/arxiv.2507.06513
openalex publication_date 2025/07/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Advances in vision-based sensors and computer vision algorithms have significantly improved the analysis and understanding of traffic scenarios. To facilitate the use of these improvements for road safety, this survey systematically categorizes the critical elements that demand attention in traffic scenarios and comprehensively analyzes available vision-driven tasks and datasets. Compared to existing surveys that focus on isolated domains, our taxonomy categorizes attention-worthy traffic entities into two main groups that are anomalies and normal but critical entities, integrating ten categories and twenty subclasses. It establishes connections between inherently related fields and provides a unified analytical framework. Our survey highlights the analysis of 35 vision-driven tasks and comprehensive examinations and visualizations of 73 available datasets based on the proposed taxonomy. The cross-domain investigation covers the pros and cons of each benchmark with the aim of providing information on standards unification and resource optimization. Our article concludes with a systematic discussion of the existing weaknesses, underlining the potential effects and promising solutions from various perspectives. The integrated taxonomy, comprehensive analysis, and recapitulatory tables serve as valuable contributions to this rapidly evolving field by providing researchers with a holistic overview, guiding strategic resource selection, and highlighting critical research gaps.
Citations
- Interaction-Centric Knowledge Infusion and Transfer for Open-Vocabulary Scene Graph Generation
- Waymo-3DSkelMo: A Multi-Agent 3D Skeletal Motion Dataset for Pedestrian Interaction Modeling in Autonomous Driving
- OccCylindrical: Multi-Modal Fusion with Cylindrical Representation for 3D Semantic Occupancy Prediction
- Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous Driving
- M2S-RoAD: Multi-Modal Semantic Segmentation for Road Damage Using Camera and LiDAR Data
- 3D Occupancy Prediction with Low-Resolution Queries via Prototype-aware View Transformation
- Motion Forecasting for Autonomous Vehicles: A Survey
- IS-Fusion: Instance-Scene Collaborative Fusion for Multimodal 3D Object Detection
- Unleashing HyDRa: Hybrid Fusion, Depth Consistency and Radar for Unified 3D Perception
- OccFusion: Multi-Sensor Fusion Framework for 3D Semantic Occupancy Prediction
- Abductive Ego-View Accident Video Understanding for Safe Driving Perception
- Vehicle Behavior Prediction by Episodic-Memory Implanted NDT
- Advancing Video Anomaly Detection: A Concise Review and a New Dataset
- Road Surface Defect Detection -- From Image-based to Non-image-based: A Survey
- A Survey on Autonomous Driving Datasets: Statistics, Annotation Quality, and a Future Outlook
- Online Vectorized HD Map Construction using Geometry
- RiskBench: A Scenario-based Benchmark for Risk Identification
- Vision-Based Traffic Accident Detection and Anticipation: A Survey
- Recursive Video Lane Detection
- Survey on video anomaly detection in dynamic scenes with moving cameras
- UGainS: Uncertainty Guided Anomaly Instance Segmentation
- Scene as Occupancy
- Weakly Supervised Multi-Modal 3D Human Body Pose Estimation for Autonomous Driving
- Occ3D: A Large-Scale 3D Occupancy Prediction Benchmark for Autonomous Driving
- OpenLane-V2: A Topology Reasoning Benchmark for Unified 3D HD Mapping
- SurroundOcc: Multi-Camera 3D Occupancy Prediction for Autonomous Driving
- Trajectory-Prediction with Vision: A Survey
- FastInst: A Simple Query-Based Model for Real-Time Instance Segmentation
- OpenOccupancy: A Large Scale Benchmark for Surrounding Semantic Occupancy Perception
- Tri-Perspective View for Vision-Based 3D Semantic Occupancy Prediction
- Crowdsensing-based Road Damage Detection Challenge (CRDDC-2022)
- OneFormer: One Transformer to Rule Universal Image Segmentation
- GLARE: A Dataset for Traffic Sign Detection in Sun Glare
- RDD2022: A multi-national image dataset for automatic Road Damage Detection
- Lane Change Classification and Prediction with Action Recognition Networks
- ViP3D: End-to-end Visual Trajectory Prediction via 3D Agent Queries
- VectorMapNet: End-to-end Vectorized HD Map Learning
- Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation
- SHREC 2022: pothole and crack detection in the road pavement using images and RGB-D data
- NHA12D: A New Pavement Crack Dataset and a Comparison Study Of Crack Detection Algorithms
- PersFormer: 3D Lane Detection via Perspective Transformer and the OpenLane Benchmark
- RelTR: Relation Transformer for Scene Graph Generation
- Masked-attention Mask Transformer for Universal Image Segmentation
- MonoScene: Monocular 3D Semantic Scene Completion
- Spatio-Temporal Scene-Graph Embedding for Autonomous Vehicle Collision\n Prediction
- LLVIP: A Visible-infrared Paired Dataset for Low-light Vision
- VIL-100: A New Dataset and A Baseline Model for Video Instance Lane Detection
- DRIVE: Deep Reinforced Accident Anticipation with Visual Explanation
- SOLQ: Segmenting Objects by Learning Queries
- SegmentMeIfYouCan: A Benchmark for Anomaly Segmentation
- Bipartite Graph Network with Adaptive Message Passing for Unbiased Scene Graph Generation
- Video action recognition for lane-change classification and prediction of surrounding vehicles
- Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers
- Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers
- Detecting Road Obstacles by Erasing Them
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- A Driving Behavior Recognition Model with Bi-LSTM and Multi-Scale CNN
- Driver Anomaly Detection: A Dataset and Contrastive Learning Approach
- Scene-Graph Augmented Data-Driven Risk Assessment of Autonomous Vehicle Decisions
- Evaluation for Weakly Supervised Object Localization: Protocol, Metrics, and Datasets
- Center-based 3D Object Detection and Tracking
- Map-Guided Curriculum Domain Adaptation and Uncertainty-Aware Evaluation for Semantic Nighttime Image Segmentation
- An Iteratively Optimized Patch Label Inference Network for Automatic Pavement Distress Detection
- End-to-End Object Detection with Transformers
- End-to-End Lane Marker Detection via Row-wise Classification
- SCRDet++: Detecting Small, Cluttered and Rotated Objects via Instance-Level Feature Denoising and Rotation Loss Smoothing
- When, Where, and What? A New Dataset for Anomaly Detection in Driving Videos
- Spatiotemporal Relationship Reasoning for Pedestrian Intent Prediction
- Bridging Knowledge Graphs to Generate Scene Graphs
- DADA: Driver Attention Prediction in Driving Accident Scenarios
- PointRend: Image Segmentation as Rendering
- SOLO: Segmenting Objects by Locations
- Scalability in Perception for Autonomous Driving: Waymo Open Dataset
- Scalability in Perception for Autonomous Driving: Waymo Open Dataset
- The inD Dataset: A Drone Dataset of Naturalistic Road User Trajectories at German Intersections
- Panoptic-DeepLab: A Simple, Strong, and Fast Baseline for Bottom-Up Panoptic Segmentation
- DADA-2000: Can Driving Accident be Predicted by Driver Attention? Analyzed by A Benchmark
- Detecting the Unexpected via Image Resynthesis
- The Fishyscapes Benchmark: Measuring Blind Spots in Semantic Segmentation
- FCOS: Fully Convolutional One-Stage Object Detection
- Deep Learning for Large-Scale Traffic-Sign Detection and Recognition
- nuScenes: A multimodal dataset for autonomous driving
- Knowledge-Embedded Routing Network for Scene Graph Generation
- Unsupervised Traffic Accident Detection in First-Person Videos
- DeeperLab: Single-Shot Image Parser
- Feature Pyramid and Hierarchical Boosting Network for Pavement Crack Detection
- Panoptic Feature Pyramid Networks
- OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields
- OpenPose: Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields
- 3D-LaneNet: End-to-End 3D Multiple Lane Detection
- Dark Model Adaptation: Semantic Image Segmentation from Daytime to Nighttime
- Egocentric Vision-based Future Vehicle Localization for Intelligent\n Driving Assistance Systems
- CADP: A Novel Dataset for CCTV Traffic Camera based Accident Analysis
- PedX: Benchmark Dataset for Metric 3D Pose Estimation of Pedestrians in Complex Urban Intersections
- Model Adaptation with Synthetic and Real Data for Semantic Dense Foggy\n Scene Understanding
- BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning
- BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning
- DensePose: Dense Human Pose Estimation In The Wild
- Real-world Anomaly Detection in Surveillance Videos
- Panoptic Segmentation
- Spatial As Deep: Spatial CNN for Traffic Scene Understanding
- Predicting Driver Attention in Critical Situations
- CARLA: An Open Urban Driving Simulator
- Semantic Foggy Scene Understanding with Synthetic Data
- Predicting the Driver's Focus of Attention: the DR(eye)VE Project
- Mask R-CNN
- CityPersons: A Diverse Dataset for Pedestrian Detection
- Scene Graph Generation by Iterative Message Passing
- Fast On-Line Kernel Density Estimation for Active Object Localization
- Lost and Found: Detecting Small Road Hazards for Self-Driving Vehicles
- Visual Relationship Detection with Language Priors
- End-to-end training of object class detectors for mean average precision
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- The Cityscapes Dataset for Semantic Urban Scene Understanding
- Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
- Deep Residual Learning for Image Recognition
- UA-DETRAC: A New Benchmark and Protocol for Multi-Object Detection and Tracking
- You Only Look Once: Unified, Real-Time Object Detection
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal\n Networks
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- U-Net: Convolutional Networks for Biomedical Image Segmentation
- Fast R-CNN
- Fully Convolutional Networks for Semantic Segmentation
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- ImageNet Large Scale Visual Recognition Challenge
- ImageNet Large Scale Visual Recognition Challenge
- Simultaneous Detection and Segmentation
- Microsoft COCO: Common Objects in Context
Related