English
Related papers

Related papers: UniDriveDreamer: A Single-Stage Multimodal World M…

200 papers

Autonomous driving technology has advanced significantly, yet detecting driving anomalies remains a major challenge due to the long-tailed distribution of driving events. Existing methods primarily rely on single-modal road condition video…

Computer Vision and Pattern Recognition · Computer Science 2025-02-06 Long Zhouxiang , Ovanes Petrosian

World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent video generation models have demonstrated impressive visual…

End-to-End autonomous driving (E2E-AD) has emerged as a new paradigm, where trajectory planning plays a crucial role. Existing studies mainly follow two directions: trajectory generation oriented, which focuses on producing high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Bin Sun , Yaoguang Cao , Yan Wang , Rui Wang , Jiachen Shang , Xiejie Feng , Jiayi Lu , Jia Shi , Shichun Yang , Xiaoyu Yan , Ziying Song

Vision-Language-Action (VLA) models are emerging as a promising paradigm for end-to-end autonomous driving, valued for their potential to leverage world knowledge and reason about complex driving scenes. However, existing methods suffer…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Xinyang Wang , Qian Liu , Wenjie Ding , Zhao Yang , Wei Li , Chang Liu , Bailin Li , Kun Zhan , Xianpeng Lang , Wei Chen

The development of computer vision algorithms for Unmanned Aerial Vehicles (UAVs) imagery heavily relies on the availability of annotated high-resolution aerial data. However, the scarcity of large-scale real datasets with pixel-level…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Giulia Rizzoli , Francesco Barbato , Matteo Caligiuri , Pietro Zanuttigh

LiDAR scene generation is increasingly important for scalable simulation and synthetic data creation, especially under diverse sensing conditions that are costly to capture at scale. Typically, diffusion-based LiDAR generators are developed…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Youquan Liu , Weidong Yang , Ao Liang , Xiang Xu , Lingdong Kong , Yang Wu , Dekai Zhu , Xin Li , Runnan Chen , Ben Fei , Tongliang Liu , Wanli Ouyang

To meet the requirements for managing unauthorized UAVs in the low-altitude economy, a multi-modal UAV trajectory prediction method based on the fusion of LiDAR and millimeter-wave radar information is proposed. A deep fusion network for…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Yuan Gao , Xinyu Guo , Wenjing Xie , Zifan Wang , Hongwen Yu , Gongyang Li , Shugong Xu

This paper proposes a unified diffusion framework (dubbed UniDiffuser) to fit all distributions relevant to a set of multi-modal data in one model. Our key insight is -- learning diffusion models for marginal, conditional, and joint…

Machine Learning · Computer Science 2023-05-31 Fan Bao , Shen Nie , Kaiwen Xue , Chongxuan Li , Shi Pu , Yaole Wang , Gang Yue , Yue Cao , Hang Su , Jun Zhu

Recent video generation models largely rely on video autoencoders that compress pixel-space videos into latent representations. However, existing video autoencoders suffer from three major limitations: (1) fixed-rate compression that wastes…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Yao Teng , Minxuan Lin , Xian Liu , Shuai Wang , Xiao Yang , Xihui Liu

Learning a robust video Variational Autoencoder (VAE) is essential for reducing video redundancy and facilitating efficient video generation. Directly applying image VAEs to individual frames in isolation can result in temporal…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Yazhou Xing , Yang Fei , Yingqing He , Jingye Chen , Jiaxin Xie , Xiaowei Chi , Qifeng Chen

Driving world models are used to simulate futures by video generation based on the condition of the current state and actions. However, current models often suffer serious error accumulations when predicting the long-term future, which…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Xiaodong Wang , Zhirong Wu , Peixi Peng

Scaling Vision-Language-Action (VLA) models on large-scale data offers a promising path to achieving a more generalized driving intelligence. However, VLA models are limited by a ``supervision deficit'': the vast model capacity is…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Yingyan Li , Shuyao Shang , Weisong Liu , Bing Zhan , Haochen Wang , Yuqi Wang , Yuntao Chen , Xiaoman Wang , Yasong An , Chufeng Tang , Lu Hou , Lue Fan , Zhaoxiang Zhang

Modeling dynamic 3D environments from LiDAR sequences is central to building reliable 4D worlds for autonomous driving and embodied AI. Existing generative frameworks, however, often treat all spatial regions uniformly, overlooking the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Xiang Xu , Alan Liang , Youquan Liu , Linfeng Li , Lingdong Kong , Ziwei Liu , Qingshan Liu

Camera control, which achieves diverse visual effects by changing camera position and pose, has attracted widespread attention. However, existing methods face challenges such as complex interaction and limited control capabilities. To…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Xiaoda Yang , Jiayang Xu , Kaixuan Luan , Xinyu Zhan , Hongshun Qiu , Shijun Shi , Hao Li , Shuai Yang , Li Zhang , Checheng Yu , Cewu Lu , Lixin Yang

End-to-end autonomous driving (E2E-AD) has emerged as a promising paradigm that unifies perception, prediction, and planning into a holistic, data-driven framework. However, achieving robustness to varying camera viewpoints, a common…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Hoonhee Cho , Jae-Young Kang , Giwon Lee , Hyemin Yang , Heejun Park , Seokwoo Jung , Kuk-Jin Yoon

In this paper, a multi-modal 360$^{\circ}$ framework for 3D object detection and tracking for autonomous vehicles is presented. The process is divided into four main stages. First, images are fed into a CNN network to obtain instance…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Jorge Beltrán , Carlos Guindel , Irene Cortés , Alejandro Barrera , Armando Astudillo , Jesús Urdiales , Mario Álvarez , Farid Bekka , Vicente Milanés , Fernando García

We propose the use of latent space generative world models to address the covariate shift problem in autonomous driving. A world model is a neural network capable of predicting an agent's next state given past states and actions. By…

Medical diagnostic applications require models that can process multimodal medical inputs (images, patient histories, lab results) and generate diverse outputs including both textual reports and visual content (annotations, segmentation…

Accurate accident anticipation remains challenging when driver cognition and dynamic road conditions are underrepresented in predictive models. In this paper, we propose CAMERA (Context-Aware Multi-modal Enhanced Risk Anticipation), a…

Computational Engineering, Finance, and Science · Computer Science 2025-07-17 Jiaxun Zhang , Haicheng Liao , Yumu Xie , Chengyue Wang , Yanchen Guan , Bin Rao , Zhenning Li

World models and video generation are pivotal technologies in the domain of autonomous driving, each playing a critical role in enhancing the robustness and reliability of autonomous systems. World models, which simulate the dynamics of…

Artificial Intelligence · Computer Science 2024-11-06 Ao Fu , Yi Zhou , Tao Zhou , Yi Yang , Bojun Gao , Qun Li , Guobin Wu , Ling Shao
‹ Prev 1 8 9 10 Next ›