English
Related papers

Related papers: Temporal-Channel Transformer for 3D Lidar-Based Vi…

200 papers

This paper presents Camera-LiDAR Fusion Transformer (CLFT) models for traffic object segmentation, which leverage the fusion of camera and LiDAR data using vision transformers. Building on the methodology of visual transformers that exploit…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Toomas Tahves , Junyi Gu , Mauro Bellone , Raivo Sell

Query-based transformer has shown great potential in constructing long-range attention in many image-domain tasks, but has rarely been considered in LiDAR-based 3D object detection due to the overwhelming size of the point cloud data. In…

Computer Vision and Pattern Recognition · Computer Science 2022-09-14 Zixiang Zhou , Xiangchen Zhao , Yu Wang , Panqu Wang , Hassan Foroosh

Vision-based Transformer have shown huge application in the perception module of autonomous driving in terms of predicting accurate 3D bounding boxes, owing to their strong capability in modeling long-range dependencies between the visual…

Computer Vision and Pattern Recognition · Computer Science 2023-04-06 Apoorv Singh

Detecting abnormal activities in real-world surveillance videos is an important yet challenging task as the prior knowledge about video anomalies is usually limited or unavailable. Despite that many approaches have been developed to resolve…

Computer Vision and Pattern Recognition · Computer Science 2021-07-30 Xinyang Feng , Dongjin Song , Yuncong Chen , Zhengzhang Chen , Jingchao Ni , Haifeng Chen

In autonomous driving, accurate 3D lane detection using monocular cameras is important for downstream tasks. Recent CNN and Transformer approaches usually apply a two-stage model design. The first stage transforms the image feature from a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Yifeng Bai , Zhirong Chen , Pengpeng Liang , Bo Song , Erkang Cheng

We describe a new spatio-temporal video autoencoder, based on a classic spatial image autoencoder and a novel nested temporal autoencoder. The temporal encoder is represented by a differentiable visual memory composed of convolutional long…

Machine Learning · Computer Science 2016-09-02 Viorica Patraucean , Ankur Handa , Roberto Cipolla

We consider the problem of localizing a spatio-temporal tube in a video corresponding to a given text query. This is a challenging task that requires the joint and efficient modeling of temporal, spatial and multi-modal interactions. To…

Computer Vision and Pattern Recognition · Computer Science 2022-06-10 Antoine Yang , Antoine Miech , Josef Sivic , Ivan Laptev , Cordelia Schmid

Camera and radar sensors have significant advantages in cost, reliability, and maintenance compared to LiDAR. Existing fusion methods often fuse the outputs of single modalities at the result-level, called the late fusion strategy. This can…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Youngseok Kim , Sanmin Kim , Jun Won Choi , Dongsuk Kum

The task of retrieving video content relevant to natural language queries plays a critical role in effectively handling internet-scale datasets. Most of the existing methods for this caption-to-video retrieval problem do not fully exploit…

Computer Vision and Pattern Recognition · Computer Science 2020-07-22 Valentin Gabeur , Chen Sun , Karteek Alahari , Cordelia Schmid

Estimating and understanding the surroundings of the vehicle precisely forms the basic and crucial step for the autonomous vehicle. The perception system plays a significant role in providing an accurate interpretation of a vehicle's…

Computer Vision and Pattern Recognition · Computer Science 2022-03-16 Sreenivasa Hikkal Venugopala

Accurate and robust LiDAR 3D object detection is essential for comprehensive scene understanding in autonomous driving. Despite its importance, LiDAR detection performance is limited by inherent constraints of point cloud data, particularly…

Computer Vision and Pattern Recognition · Computer Science 2024-09-09 Rui Yu , Runkai Zhao , Cong Nie , Heng Wang , HuaiCheng Yan , Meng Wang

Convolutional neural networks have made significant progresses in edge detection by progressively exploring the context and semantic features. However, local details are gradually suppressed with the enlarging of receptive fields. Recently,…

Computer Vision and Pattern Recognition · Computer Science 2022-03-17 Mengyang Pu , Yaping Huang , Yuming Liu , Qingji Guan , Haibin Ling

Urban-oriented autonomous vehicles require a reliable perception technology to tackle the high amount of uncertainties. The recently introduced compact 3D LIDAR sensor offers a surround spatial information that can be exploited to enhance…

Computer Vision and Pattern Recognition · Computer Science 2018-04-24 Achim Kampker , Mohsen Sefati , Arya Abdul Rachman , Kai Kreisköther , Pascual Campoy

Understanding natural-language references to objects in dynamic 3D driving scenes is essential for interactive autonomous systems. In practice, many referring expressions describe targets through recent motion or short-term interactions,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Jiahong Yu , Ziqi Wang , Hailiang Zhao , Wei Zhai , Xueqiang Yan , Shuiguang Deng

Multi-sensor fusion using LiDAR and RGB cameras significantly enhances 3D object detection task. However, conventional LiDAR sensors perform dense, stateless scans, ignoring the strong temporal continuity in real-world scenes. This leads to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Sara Shoouri , Morteza Tavakoli Taba , Hun-Seok Kim

Human detection and tracking is an essential task for service robots, where the combined use of multiple sensors has potential advantages that are yet to be exploited. In this paper, we introduce a framework allowing a robot to learn a new…

Robotics · Computer Science 2018-08-01 Zhi Yan , Li Sun , Tom Duckett , Nicola Bellotto

Multi-level features are important for saliency detection. Better combination and use of multi-level features with time information can greatly improve the accuracy of the video saliency model. In order to fully combine multi-level features…

Computer Vision and Pattern Recognition · Computer Science 2021-09-15 Qinyao Chang , Shiping Zhu

There has been significant progress made in the field of autonomous vehicles. Object detection and tracking are the primary tasks for any autonomous vehicle. The task of object detection in autonomous vehicles relies on a variety of sensors…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Gaurav Raut , Advait Patole

Fully autonomous driving systems require fast detection and recognition of sensitive objects in the environment. In this context, intelligent vehicles should share their sensor data with computing platforms and/or other vehicles, to detect…

Networking and Internet Architecture · Computer Science 2021-04-27 Valentina Rossi , Paolo Testolina , Marco Giordani , Michele Zorzi

This paper presents a novel 3D object detection framework that processes LiDAR data directly on its native representation: range images. Benefiting from the compactness of range images, 2D convolutions can efficiently process dense LiDAR…

Computer Vision and Pattern Recognition · Computer Science 2021-01-25 Alex Bewley , Pei Sun , Thomas Mensink , Dragomir Anguelov , Cristian Sminchisescu