中文
相关论文

相关论文: VisionPAD: A Vision-Centric Pre-training Paradigm …

200 篇论文

Although existing monocular depth estimation methods have made great progress, predicting an accurate absolute depth map from a single image is still challenging due to the limited modeling capacity of networks and the scale ambiguity…

计算机视觉与模式识别 · 计算机科学 2022-10-07 Jie Xiang , Yun Wang , Lifeng An , Haiyang Liu , Zijun Wang , Jian Liu

Robotic grasping is a fundamental capability for enabling autonomous manipulation, with usually infinite solutions. State-of-the-art approaches for grasping rely on learning from large-scale datasets comprising expert annotations of…

机器人学 · 计算机科学 2026-03-17 Manav Kulshrestha , S. Talha Bukhari , Damon Conover , Aniket Bera

Realistic scene reconstruction and view synthesis are essential for advancing autonomous driving systems by simulating safety-critical scenarios. 3D Gaussian Splatting excels in real-time rendering and static scene reconstructions but…

计算机视觉与模式识别 · 计算机科学 2024-07-08 Mustafa Khan , Hamidreza Fazlali , Dhruv Sharma , Tongtong Cao , Dongfeng Bai , Yuan Ren , Bingbing Liu

The significant achievements of pre-trained models leveraging large volumes of data in the field of NLP and 2D vision inspire us to explore the potential of extensive data pre-training for 3D perception in autonomous driving. Toward this…

计算机视觉与模式识别 · 计算机科学 2025-04-18 Shumin Wang , Zhuoran Yang , Lidian Wang , Zhipeng Tang , Heng Li , Lehan Pan , Sha Zhang , Jie Peng , Jianmin Ji , Yanyong Zhang

Vision-based deep learning (DL) methods have made great progress in learning autonomous driving models from large-scale crowd-sourced video datasets. They are trained to predict instantaneous driving behaviors from video data captured by…

人机交互 · 计算机科学 2021-09-24 Suphanut Jamonnak , Ye Zhao , Xinyi Huang , Md Amiruzzaman

Humans excel at efficiently navigating through crowds without collision by focusing on specific visual regions relevant to navigation. However, most robotic visual navigation methods rely on deep learning models pre-trained on vision tasks,…

机器人学 · 计算机科学 2025-01-03 Mohammad Nazeri , Junzhe Wang , Amirreza Payandeh , Xuesu Xiao

In this paper we present a novel unsupervised representation learning approach for 3D shapes, which is an important research challenge as it avoids the manual effort required for collecting supervised data. Our method trains an RNN-based…

计算机视觉与模式识别 · 计算机科学 2018-11-08 Zhizhong Han , Mingyang Shang , Yu-Shen Liu , Matthias Zwicker

Autonomous driving is a challenging task that requires perceiving and understanding the surrounding environment for safe trajectory planning. While existing vision-based end-to-end models have achieved promising results, these methods are…

计算机视觉与模式识别 · 计算机科学 2025-01-16 Tengpeng Li , Hanli Wang , Xianfei Li , Wenlong Liao , Tao He , Pai Peng

End-to-end autonomous driving systems promise stronger performance through unified optimization of perception, motion forecasting, and planning. However, vision-based approaches face fundamental limitations in adverse weather conditions,…

计算机视觉与模式识别 · 计算机科学 2025-08-15 Philipp Wolters , Johannes Gilg , Torben Teepe , Gerhard Rigoll

3D object detection and dense depth estimation are one of the most vital tasks in autonomous driving. Multiple sensor modalities can jointly attribute towards better robot perception, and to that end, we introduce a method for jointly…

计算机视觉与模式识别 · 计算机科学 2021-09-16 Shubham Shrivastava

The perception of autonomous vehicles using radars has attracted increased research interest due its ability to operate in fog and bad weather. However, training radar models is hindered by the cost and difficulty of annotating large-scale…

计算机视觉与模式识别 · 计算机科学 2024-04-19 Yiduo Hao , Sohrab Madani , Junfeng Guan , Mohammed Alloulah , Saurabh Gupta , Haitham Hassanieh

High-performing vision language models still produce incorrect answers, yet their failure modes are often difficult to explain. To make model internals more accessible and enable systematic debugging, we introduce VisualScratchpad, an…

人工智能 · 计算机科学 2026-03-10 Hyesu Lim , Jinho Choi , Taekyung Kim , Byeongho Heo , Jaegul Choo , Dongyoon Han

This paper targets on learning-based novel view synthesis from a single or limited 2D images without the pose supervision. In the viewer-centered coordinates, we construct an end-to-end trainable conditional variational framework to…

计算机视觉与模式识别 · 计算机科学 2021-06-08 Xiaofeng Liu , Tong Che , Yiqun Lu , Chao Yang , Site Li , Jane You

Learning robust and scalable visual representations from massive multi-view video data remains a challenge in computer vision and autonomous driving. Existing pre-training methods either rely on expensive supervised learning with 3D…

计算机视觉与模式识别 · 计算机科学 2024-03-14 Jialv Zou , Bencheng Liao , Qian Zhang , Wenyu Liu , Xinggang Wang

In recent years, vision-based end-to-end autonomous driving has emerged as a new paradigm. However, popular end-to-end approaches typically rely on visual feature extraction networks trained under label supervision. This limited supervision…

机器人学 · 计算机科学 2025-11-04 Ling Niu , Xiaoji Zheng , Han Wang , Chen Zheng , Ziyuan Yang , Bokui Chen , Jiangtao Gong

Recent advancements in camera-based occupancy prediction have focused on the simultaneous prediction of 3D semantics and scene flow, a task that presents significant challenges due to specific difficulties, e.g., occlusions and unbalanced…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Ziyue Zhu , Shenlong Wang , Jin Xie , Jiang-jiang Liu , Jingdong Wang , Jian Yang

Vision-based autonomous urban driving in dense traffic is quite challenging due to the complicated urban environment and the dynamics of the driving behaviors. Widely-applied methods either heavily rely on hand-crafted rules or learn from…

计算机视觉与模式识别 · 计算机科学 2023-04-20 Yinuo Zhao , Kun Wu , Zhiyuan Xu , Zhengping Che , Qi Lu , Jian Tang , Chi Harold Liu

Autonomous driving faces safety challenges due to a lack of global perspective and the semantic information of vectorized high-definition (HD) maps. Information from roadside cameras can greatly expand the map perception range through…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Miao Fan , Shanshan Yu , Shengtong Xu , Kun Jiang , Haoyi Xiong , Xiangzeng Liu

Understanding the 3D world without supervision is currently a major challenge in computer vision as the annotations required to supervise deep networks for tasks in this domain are expensive to obtain on a large scale. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2022-12-02 Octave Mariotti , Oisin Mac Aodha , Hakan Bilen

We present an integrated approach for perception and control for an autonomous vehicle and demonstrate this approach in a high-fidelity urban driving simulator. Our approach first builds a model for the environment, then trains a policy…

系统与控制 · 电气工程与系统科学 2020-03-19 Ali Baheri , Ilya Kolmanovsky , Anouck Girard , H. Eric Tseng , Dimitar Filev