中文
相关论文

相关论文: VisionPAD: A Vision-Centric Pre-training Paradigm …

200 篇论文

Feed-forward 3D Gaussian Splatting (3DGS) has emerged as a highly effective solution for novel view synthesis. Existing methods predominantly rely on a \emph{pixel-aligned} Gaussian prediction paradigm, where each 2D pixel is mapped to a 3D…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Weijie Wang , Yeqing Chen , Zeyu Zhang , Hengyu Liu , Haoxiao Wang , Zhiyuan Feng , Wenkang Qin , Feng Chen , Zheng Zhu , Donny Y. Chen , Bohan Zhuang

Accurately understanding and deciding high-level meta-actions is essential for ensuring reliable and safe autonomous driving systems. While vision-language models (VLMs) have shown significant potential in various autonomous driving tasks,…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Yujin Wang , Quanfeng Liu , Zhengxin Jiang , Tianyi Wang , Junfeng Jiao , Hongqing Chu , Bingzhao Gao , Hong Chen

Accurate 3D perception is essential for autonomous driving. Traditional methods often struggle with geometric ambiguity due to a lack of geometric prior. To address these challenges, we use omnidirectional depth estimation to introduce…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Chaofan Wu , Jiaheng Li , Jinghao Cao , Ming Li , Yongkang Feng , Jiayu Wu Shuwen Xu , Zihang Gao , Sidan Du , Yang Li

Self-supervised pre-training based on next-token prediction has enabled large language models to capture the underlying structure of text, and has led to unprecedented performance on a large array of tasks when applied at scale. Similarly,…

Face Presentation Attack Detection (PAD) is an important measure to prevent spoof attacks for face biometric systems. Many works based on Convolution Neural Networks (CNNs) for face PAD formulate the problem as an image-level binary…

计算机视觉与模式识别 · 计算机科学 2022-03-15 Zuheng Ming , Zitong Yu , Musab Al-Ghadi , Muriel Visani , Muhammad MuzzamilLuqman , Jean-Christophe Burie

Autonomous driving is a complex and challenging task that aims at safe motion planning through scene understanding and reasoning. While vision-only autonomous driving methods have recently achieved notable performance, through enhanced…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Chenbin Pan , Burhaneddin Yaman , Tommaso Nesti , Abhirup Mallik , Alessandro G Allievi , Senem Velipasalar , Liu Ren

The goal of our work is to use visual attention to enhance autonomous driving performance. We present two methods of predicting visual attention maps. The first method is a supervised learning approach in which we collect eye-gaze data for…

计算机视觉与模式识别 · 计算机科学 2018-12-06 Sourav Pal , Tharun Mohandoss , Pabitra Mitra

Vision-based learning methods for self-driving cars have primarily used supervised approaches that require a large number of labels for training. However, those labels are usually difficult and expensive to obtain. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2021-03-23 Qadeer Khan , Patrick Wenzel , Daniel Cremers

Perception of the environment is a critical component for enabling autonomous driving. It provides the vehicle with the ability to comprehend its surroundings and make informed decisions. Depth prediction plays a pivotal role in this…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Houssem Boulahbal

Vision-based roadside 3D object detection has attracted rising attention in autonomous driving domain, since it encompasses inherent advantages in reducing blind spots and expanding perception range. While previous work mainly focuses on…

计算机视觉与模式识别 · 计算机科学 2024-06-14 Wenjie Wang , Yehao Lu , Guangcong Zheng , Shuigen Zhan , Xiaoqing Ye , Zichang Tan , Jingdong Wang , Gaoang Wang , Xi Li

Recent works have shown that visual pretraining on egocentric datasets using masked autoencoders (MAE) can improve generalization for downstream robotics tasks. However, these approaches pretrain only on 2D images, while many robotics…

机器人学 · 计算机科学 2025-03-25 Shengyi Qian , Kaichun Mo , Valts Blukis , David F. Fouhey , Dieter Fox , Ankit Goyal

Dynamic Gaussian splatting has led to impressive scene reconstruction and image synthesis advances in novel views. Existing methods, however, heavily rely on pre-computed poses and Gaussian initialization by Structure from Motion (SfM)…

计算机视觉与模式识别 · 计算机科学 2024-06-27 Hao Li , Jingfeng Li , Dingwen Zhang , Chenming Wu , Jieqi Shi , Chen Zhao , Haocheng Feng , Errui Ding , Jingdong Wang , Junwei Han

Large vision-language models (VLMs) have shown promising capabilities in scene understanding, enhancing the explainability of driving behaviors and interactivity with users. Existing methods primarily fine-tune VLMs on on-board multi-view…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Nan Song , Bozhou Zhang , Xiatian Zhu , Jiankang Deng , Li Zhang

Surround depth estimation provides a cost-effective alternative to LiDAR for 3D perception in autonomous driving. While recent self-supervised methods explore multi-camera settings to improve scale awareness and scene coverage, they are…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Weimin Liu , Jiyuan Qiu , Wenjun Wang , Joshua H. Meng

This dissertation is a multifaceted contribution to the advancement of vision-based 3D perception technologies. In the first segment, the thesis introduces structural enhancements to both monocular and stereo 3D object detection algorithms.…

计算机视觉与模式识别 · 计算机科学 2024-03-05 Yuxuan Liu

We present ViewSplat, a view-adaptive 3D Gaussian splatting network for novel view synthesis from unposed images. While recent feed-forward 3D Gaussian splatting has significantly accelerated 3D scene reconstruction by bypassing per-scene…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Moonyeon Jeong , Seunggi Min , Suhyeon Lee , Hongje Seong

Existing multi-view image compression methods often rely on 2D projection-based similarities between views to estimate disparities. While effective for small disparities, such as those in stereo images, these methods struggle with the more…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Yujun Huang , Bin Chen , Niu Lian , Baoyi An , Shu-Tao Xia

Robust trajectory planning under camera viewpoint changes is important for scalable end-to-end autonomous driving. However, existing models often depend heavily on the camera viewpoints seen during training. We investigate an…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Hiroki Hashimoto , Hiromichi Goto , Hiroyuki Sugai , Hiroshi Kera , Kazuhiko Kawamoto

In the realm of autonomous driving, accurately detecting surrounding obstacles is crucial for effective decision-making. Traditional methods primarily rely on 3D bounding boxes to represent these obstacles, which often fail to capture the…

机器人学 · 计算机科学 2025-11-18 Chunyong Hu , Qi Luo , Jianyun Xu , Song Wang , Qiang Li , Sheng Yang

We propose a new self-supervised method for pre-training the backbone of deep perception models operating on point clouds. The core idea is to train the model on a pretext task which is the reconstruction of the surface on which the 3D…

计算机视觉与模式识别 · 计算机科学 2023-04-05 Alexandre Boulch , Corentin Sautier , Björn Michele , Gilles Puy , Renaud Marlet