English
Related papers

Related papers: ET-Former: Efficient Triplane Deformable Attention…

200 papers

Accurate perception of the dynamic environment is a fundamental task for autonomous driving and robot systems. This paper introduces Let Occ Flow, the first self-supervised work for joint 3D occupancy and occupancy flow prediction using…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Yili Liu , Linzhan Mou , Xuan Yu , Chenrui Han , Sitong Mao , Rong Xiong , Yue Wang

Camera-based 3D semantic scene completion (SSC) provides dense geometric and semantic perception for autonomous driving and robotic navigation. However, existing methods rely on a coupled encoder to deliver both semantic and geometric…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Shiyuan Chen , Wei Sui , Bohao Zhang , Zeyd Boukhers , John See , Cong Yang

We present SLAM-Former, a novel neural approach that integrates full SLAM capabilities into a single transformer. Similar to traditional SLAM systems, SLAM-Former comprises both a frontend and a backend that operate in tandem. The frontend…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Yijun Yuan , Zhuoguang Chen , Kenan Li , Weibang Wang , Hang Zhao

Vision Transformer has demonstrated impressive success across various vision tasks. However, its heavy computation cost, which grows quadratically with respect to the token sequence length, largely limits its power in handling large feature…

Computer Vision and Pattern Recognition · Computer Science 2023-08-24 Sucheng Ren , Xingyi Yang , Songhua Liu , Xinchao Wang

Scene analysis is essential for enabling autonomous systems, such as mobile robots, to operate in real-world environments. However, obtaining a comprehensive understanding of the scene requires solving multiple tasks, such as panoptic…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Söhnke Benedikt Fischedick , Daniel Seichter , Robin Schmidt , Leonard Rabes , Horst-Michael Gross

Perceiving a three-dimensional (3D) scene with multiple objects while moving indoors is essential for vision-based mobile cobots, especially for enhancing their manipulation tasks. In this work, we present an end-to-end pipeline with…

Robotics · Computer Science 2024-02-20 K. Nguyen , T. Dang , M. Huber

Transparent object perception remains a major challenge in computer vision research, as transparency confounds both depth estimation and semantic segmentation. Recent work has explored multi-task learning frameworks to improve robustness,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Gbenga Omotara , Ramy Farag , Seyed Mohamad Ali Tousi , G. N. DeSouza

We present a novel framework for dynamic radiance field prediction given monocular video streams. Unlike previous methods that primarily focus on predicting future frames, our method goes a step further by generating explicit 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-01-29 Di Qi , Tong Yang , Beining Wang , Xiangyu Zhang , Wenqiang Zhang

Convolutional Neural Networks (CNNs) and Transformers have achieved remarkable success in computer vision tasks. However, their deep architectures often lead to high computational redundancy, making them less suitable for…

Computer Vision and Pattern Recognition · Computer Science 2024-10-03 Novendra Setyawan , Ghufron Wahyu Kurniawan , Chi-Chia Sun , Jun-Wei Hsieh , Jing-Ming Guo , Wen-Kai Kuo

Recovering temporally consistent 3D human body pose, shape and motion from a monocular video is a challenging task due to (self-)occlusions, poor lighting conditions, complex articulated body poses, depth ambiguity, and limited availability…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Sushovan Chanda , Amogh Tiwari , Lokender Tiwari , Brojeshwar Bhowmick , Avinash Sharma , Hrishav Barua

Autonomous navigation in marine environments can be extremely challenging, especially in the presence of spatially varying flow disturbances and dynamic and static obstacles. In this work, we demonstrate that incorporating local flow field…

Robotics · Computer Science 2025-07-11 Ehsan Kazemi , Dechen Gao , Iman Soltani

Most hard attention models initially observe a complete scene to locate and sense informative glimpses, and predict class-label of a scene based on glimpses. However, in many applications (e.g., aerial imaging), observing an entire scene is…

Computer Vision and Pattern Recognition · Computer Science 2022-04-05 Samrudhdhi B. Rangrej , Chetan L. Srinidhi , James J. Clark

3D visual perception tasks, including 3D detection and map segmentation based on multi-camera images, are essential for autonomous driving systems. In this work, we present a new framework termed BEVFormer, which learns unified BEV…

Computer Vision and Pattern Recognition · Computer Science 2022-07-14 Zhiqi Li , Wenhai Wang , Hongyang Li , Enze Xie , Chonghao Sima , Tong Lu , Qiao Yu , Jifeng Dai

Scene graph generation (SGG) of surgical procedures is crucial in enhancing holistically cognitive intelligence in the operating room (OR). However, previous works have primarily relied on multi-stage learning, where the generated semantic…

Computer Vision and Pattern Recognition · Computer Science 2024-08-08 Jialun Pei , Diandian Guo , Jingyang Zhang , Manxi Lin , Yueming Jin , Pheng-Ann Heng

Panoramic semantic segmentation models are typically trained under a strict gravity-aligned assumption. However, real-world captures often deviate from this canonical orientation due to unconstrained camera motions, such as the rotational…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Qinfeng Zhu , Yunxi Jiang , Lei Fan

For supervised speech enhancement, contextual information is important for accurate spectral mapping. However, commonly used deep neural networks (DNNs) are limited in capturing temporal contexts. To leverage long-term contexts for tracking…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-13 Xinmeng Xu , Jianjun Hao

Deep neural networks, especially transformer-based architectures, have achieved remarkable success in semantic segmentation for environmental perception. However, existing models process video frames independently, thus failing to leverage…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Serin Varghese , Kevin Ross , Fabian Hueger , Kira Maag

Transformers have been actively studied for time-series forecasting in recent years. While often showing promising results in various scenarios, traditional Transformers are not designed to fully exploit the characteristics of time-series…

Machine Learning · Computer Science 2022-06-22 Gerald Woo , Chenghao Liu , Doyen Sahoo , Akshat Kumar , Steven Hoi

In this paper, we propose a transformer-based image matting model called MatteFormer, which takes full advantage of trimap information in the transformer block. Our method first introduces a prior-token which is a global representation of…

Computer Vision and Pattern Recognition · Computer Science 2022-03-30 GyuTae Park , SungJoon Son , JaeYoung Yoo , SeHo Kim , Nojun Kwak

Vision-based 3D semantic occupancy prediction is a critical task in 3D vision that integrates volumetric 3D reconstruction with semantic understanding. Existing methods, however, often rely on modular pipelines. These modules are typically…

Computer Vision and Pattern Recognition · Computer Science 2025-09-11 Dubing Chen , Huan Zheng , Yucheng Zhou , Xianfei Li , Wenlong Liao , Tao He , Pai Peng , Jianbing Shen