vix.ing · top · new · best · stats · spec

End-to-end Contextual Perception and Prediction with Interaction\n Transformer

2020/08/13 by Lingyun Luke Li, Bin Yang, Li, Lingyun Luke +11 · 3 citations
Computer Science · Engineering · #Autonomous Vehicle Technology and Safety #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Robotics (cs.RO) #Video Surveillance and Tracking Methods

paper · pdf · doi:10.48550/arxiv.2008.05927

openalex publication_date 2020/08/13 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28

Abstract

In this paper, we tackle the problem of detecting objects in 3D and\nforecasting their future motion in the context of self-driving. Towards this\ngoal, we design a novel approach that explicitly takes into account the\ninteractions between actors. To capture their spatial-temporal dependencies, we\npropose a recurrent neural network with a novel Transformer architecture, which\nwe call the Interaction Transformer. Importantly, our model can be trained\nend-to-end, and runs in real-time. We validate our approach on two challenging\nreal-world datasets: ATG4D and nuScenes. We show that our approach can\noutperform the state-of-the-art on both datasets. In particular, we\nsignificantly improve the social compliance between the estimated future\ntrajectories, resulting in far fewer collisions between the predicted actors.\n

Cited by

Related