English

End-to-end Contextual Perception and Prediction with Interaction Transformer

Computer Vision and Pattern Recognition 2020-08-14 v1 Robotics

Abstract

In this paper, we tackle the problem of detecting objects in 3D and forecasting their future motion in the context of self-driving. Towards this goal, we design a novel approach that explicitly takes into account the interactions between actors. To capture their spatial-temporal dependencies, we propose a recurrent neural network with a novel Transformer architecture, which we call the Interaction Transformer. Importantly, our model can be trained end-to-end, and runs in real-time. We validate our approach on two challenging real-world datasets: ATG4D and nuScenes. We show that our approach can outperform the state-of-the-art on both datasets. In particular, we significantly improve the social compliance between the estimated future trajectories, resulting in far fewer collisions between the predicted actors.

Keywords

Cite

@article{arxiv.2008.05927,
  title  = {End-to-end Contextual Perception and Prediction with Interaction Transformer},
  author = {Lingyun Luke Li and Bin Yang and Ming Liang and Wenyuan Zeng and Mengye Ren and Sean Segal and Raquel Urtasun},
  journal= {arXiv preprint arXiv:2008.05927},
  year   = {2020}
}

Comments

IROS 2020

R2 v1 2026-06-23T17:50:17.123Z