English
Related papers

Related papers: DriveCamSim: Generalizable Camera Simulation via E…

200 papers

3D object detection and occupancy prediction are critical tasks in autonomous driving, attracting significant attention. Despite the potential of recent vision-based methods, they encounter challenges under adverse conditions. Thus,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Lianqing Zheng , Jianan Liu , Runwei Guan , Long Yang , Shouyi Lu , Yuanzhe Li , Xiaokai Bai , Jie Bai , Zhixiong Ma , Hui-Liang Shen , Xichan Zhu

Precise camera pose control is crucial for video generation with diffusion models. Existing methods require fine-tuning with additional datasets containing paired videos and camera pose annotations, which are both data-intensive and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Zhenghong Zhou , Jie An , Jiebo Luo

In the past few decades, autonomous driving algorithms have made significant progress in perception, planning, and control. However, evaluating individual components does not fully reflect the performance of entire systems, highlighting the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Hongyu Zhou , Longzhong Lin , Jiabao Wang , Yichong Lu , Dongfeng Bai , Bingbing Liu , Yue Wang , Andreas Geiger , Yiyi Liao

Lane detection is an essential part of the perception sub-architecture of any automated driving (AD) or advanced driver assistance system (ADAS). When focusing on low-cost, large scale products for automated driving, model-driven approaches…

Computer Vision and Pattern Recognition · Computer Science 2021-06-25 Thomas Michalke , Di Feng , Claudius Gläser , Fabian Timm

Estimating camera motion and intrinsics from casual videos is a core challenge in computer vision. Traditional bundle-adjustment based methods, such as SfM and SLAM, struggle to perform reliably on arbitrary data. Although specialized SfM…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Felix Wimbauer , Weirong Chen , Dominik Muhle , Christian Rupprecht , Daniel Cremers

Visual Question Answering (VQA) models, which fall under the category of vision-language models, conventionally execute multiple downsampling processes on image inputs to strike a balance between computational efficiency and model…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Xirui Zhou , Lianlei Shan , Xiaolin Gui

Camera control has been extensively studied in conditioned video generation; however, performing precisely altering the camera trajectories while faithfully preserving the video content remains a challenging task. The mainstream approach to…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Dong-Yu Chen , Yixin Guo , Shuojin Yang , Tai-Jiang Mu , Shi-Min Hu

This paper introduces CameraCtrl II, a framework that enables large-scale dynamic scene exploration through a camera-controlled video diffusion model. Previous camera-conditioned video generative models suffer from diminished video dynamics…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Hao He , Ceyuan Yang , Shanchuan Lin , Yinghao Xu , Meng Wei , Liangke Gui , Qi Zhao , Gordon Wetzstein , Lu Jiang , Hongsheng Li

Despite rapid progress in autonomous driving, reliable training and evaluation of driving systems remain fundamentally constrained by the lack of scalable and interactive simulation environments. Recent generative video models achieve…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yaoru Li , Federico Landi , Marco Godi , Xin Jin , Ruiju Fu , Yufei Ma , Muyang Sun , Heyu Si , Qi Guo

Autonomous mobility systems increasingly operate in dense and dynamic environments where perception occlusions, limited sensing coverage, and multi-agent interactions pose major challenges. While onboard sensors provide essential local…

Robotics · Computer Science 2026-03-18 Yufeng Yang , Minghao Ning , Keqi Shu , Aladdin Saleh , Ehsan Hashemi , Amir Khajepour

Traditional autonomous driving methods adopt a modular design, decomposing tasks into sub-tasks. In contrast, end-to-end autonomous driving directly outputs actions from raw sensor data, avoiding error accumulation. However, training an…

Robotics · Computer Science 2024-11-22 Zeyu Dong , Yimin Zhu , Yansong Li , Kevin Mahon , Yu Sun

Driving World Models (DWMs) have been developing rapidly with the advances of generative models. However, existing DWMs lack 3D scene understanding capabilities and can only generate content conditioned on input data, without the ability to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Tianchen Deng , Xuefeng Chen , Yi Chen , Qu Chen , Yuyao Xu , Lijin Yang , Le Xu , Yu Zhang , Bo Zhang , Wuxiong Huang , Hesheng Wang

A world model is an AI system that simulates how an environment evolves under actions, enabling planning through imagined futures rather than reactive perception. Current world models, however, suffer from visual conflation: the mistaken…

Artificial Intelligence · Computer Science 2026-01-23 Zhikang Chen , Tingting Zhu

Autoregressive (AR) diffusion enables streaming, interactive long-video generation by producing frames causally, yet maintaining coherence over minute-scale horizons remains challenging due to accumulated errors, motion drift, and content…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Yifei Yu , Xiaoshan Wu , Xinting Hu , Tao Hu , Yangtian Sun , Xiaoyang Lyu , Bo Wang , Lin Ma , Yuewen Ma , Zhongrui Wang , Xiaojuan Qi

Scene understanding is essential for enhancing driver safety, generating human-centric explanations for Automated Vehicle (AV) decisions, and leveraging Artificial Intelligence (AI) for retrospective driving video analysis. This study…

Computer Vision and Pattern Recognition · Computer Science 2025-01-13 Mohammed Elhenawy , Huthaifa I. Ashqar , Andry Rakotonirainy , Taqwa I. Alhadidi , Ahmed Jaber , Mohammad Abu Tami

Even though virtual testing of Autonomous Vehicles (AVs) has been well recognized as essential for safety assessment, AV simulators are still undergoing active development. One particularly challenging question is to effectively include the…

Robotics · Computer Science 2024-02-28 Andrea Piazzoni , Jim Cherian , Justin Dauwels , Lap-Pui Chau

High-definition (HD) maps provide essential semantic information of road structures for autonomous driving systems, yet current HD map construction methods require calibrated multi-camera setups and either implicit or explicit 2D-to-BEV…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Run Wang , Chaoyi Zhou , Amir Salarpour , Xi Liu , Zhi-Qi Cheng , Feng Luo , Mert D. Pesé , Siyu Huang

With increasing automation in passenger vehicles, the study of safe and smooth occupant-vehicle interaction and control transitions is key. In this study, we focus on the development of contextual, semantically meaningful representations of…

Robotics · Computer Science 2021-07-26 Akshay Rangesh , Nachiket Deo , Ross Greer , Pujitha Gunaratne , Mohan M. Trivedi

In this endeavor, we developed a comprehensive system that processes integrated visual features derived from video frames captured by a regular camera, along with depth details obtained from a point cloud scanner. This system is designed to…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Alexander Liu

Traditionally, video is structured as a sequence of discrete image frames. Recently, however, a novel video sensing paradigm has emerged which eschews video frames entirely. These "event" sensors aim to mimic the human vision system with…

Multimedia · Computer Science 2024-08-13 Andrew Freeman
‹ Prev 1 8 9 10 Next ›