中文
相关论文

相关论文: WoVoGen: World Volume-aware Diffusion for Controll…

200 篇论文

We introduce a framework that enables both multi-view character consistency and 3D camera control in video diffusion models through a novel customization data pipeline. We train the character consistency component with recorded volumetric…

Modern video generative models based on diffusion models can produce very realistic clips, but they are computationally inefficient, often requiring minutes of GPU time for just a few seconds of video. This inefficiency poses a critical…

计算机视觉与模式识别 · 计算机科学 2026-01-15 Jieying Chen , Jeffrey Hu , Joan Lasenby , Ayush Tewari

Recent advances in image and video creation, especially AI-based image synthesis, have led to the production of numerous visual scenes that exhibit a high level of abstractness and diversity. Consequently, Visual Storytelling (VST), a task…

计算与语言 · 计算机科学 2023-12-13 Shengguang Wu , Mei Yuan , Qi Su

This report presents a unified technical system addressing the two core capabilities of world models for autonomous driving: world representation and world generation. For world representation, we propose WorldRec, a feed-forward…

Realistic and diverse traffic scenarios in large quantities are crucial for the development and validation of autonomous driving systems. However, owing to numerous difficulties in the data collection process and the reliance on intensive…

机器人学 · 计算机科学 2025-10-07 Shuo Sun , Zekai Gu , Tianchen Sun , Jiawei Sun , Chengran Yuan , Yuhang Han , Dongen Li , Marcelo H. Ang

End-to-end autonomous driving aims to generate safe and plausible planning policies from raw sensor input. Driving world models have shown great potential in learning rich representations by predicting the future evolution of a driving…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Xingtai Gui , Meijie Zhang , Tianyi Yan , Wencheng Han , Jiahao Gong , Feiyang Tan , Cheng-zhong Xu , Jianbing Shen

World models aim to endow AI systems with the ability to represent, generate, and interact with dynamic environments in a coherent and temporally consistent manner. While recent video generation models have demonstrated impressive visual…

Recent breakthroughs in autonomous driving have been propelled by advances in robust world modeling, fundamentally transforming how vehicles interpret dynamic scenes and execute safe decision-making. World models have emerged as a linchpin…

机器人学 · 计算机科学 2025-09-11 Tuo Feng , Wenguan Wang , Yi Yang

Recent advances in diffusion models have revolutionized 2D and 3D content creation, yet generating photorealistic dynamic 4D scenes remains a significant challenge. Existing dynamic 4D generation methods typically rely on distilling…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Vinayak Gupta , Yunze Man , Yu-Xiong Wang

Driving world models are used to simulate futures by video generation based on the condition of the current state and actions. However, current models often suffer serious error accumulations when predicting the long-term future, which…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Xiaodong Wang , Zhirong Wu , Peixi Peng

Generating safety-critical driving scenarios is crucial for evaluating and improving autonomous driving systems, but long-tail risky situations are rarely observed in real-world data and difficult to specify through manual scenario design.…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Hongyi Lin , Wenxiu Shi , Heye Huang , Dingyi Zhuang , Song Zhang , Yang Liu , Xiaobo Qu , Jinhua Zhao

We present RadarGen, a diffusion model for synthesizing realistic automotive radar point clouds from multi-view camera imagery. RadarGen adapts efficient image-latent diffusion to the radar domain by representing radar measurements in…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Tomer Borreda , Fangqiang Ding , Sanja Fidler , Shengyu Huang , Or Litany

Recent advancements in camera-trajectory-guided image-to-video generation offer higher precision and better support for complex camera control compared to text-based approaches. However, they also introduce significant usability challenges,…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Teng Li , Guangcong Zheng , Rui Jiang , Shuigen Zhan , Tao Wu , Yehao Lu , Yining Lin , Chuanyun Deng , Yepan Xiong , Min Chen , Lin Cheng , Xi Li

This paper aims to tackle the problem of photorealistic view synthesis from vehicle sensor data. Recent advancements in neural scene representation have achieved notable success in rendering high-quality autonomous driving scenes, but the…

计算机视觉与模式识别 · 计算机科学 2025-08-27 Yunzhi Yan , Zhen Xu , Haotong Lin , Haian Jin , Haoyu Guo , Yida Wang , Kun Zhan , Xianpeng Lang , Hujun Bao , Xiaowei Zhou , Sida Peng

Autonomous driving requires an understanding of the static environment from sensor data. Learned Bird's-Eye View (BEV) encoders are commonly used to fuse multiple inputs, and a vector decoder predicts a vectorized map representation from…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Thomas Monninger , Zihan Zhang , Zhipeng Mo , Md Zafar Anwar , Steffen Staab , Sihao Ding

World models and video generation are pivotal technologies in the domain of autonomous driving, each playing a critical role in enhancing the robustness and reliability of autonomous systems. World models, which simulate the dynamics of…

人工智能 · 计算机科学 2024-11-06 Ao Fu , Yi Zhou , Tao Zhou , Yi Yang , Bojun Gao , Qun Li , Guobin Wu , Ling Shao

In autonomous driving, vision-centric 3D detection aims to identify 3D objects from images. However, high data collection costs and diverse real-world scenarios limit the scale of training data. Once distribution shifts occur between…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Hongbin Lin , Zilu Guo , Yifan Zhang , Shuaicheng Niu , Yafeng Li , Ruimao Zhang , Shuguang Cui , Zhen Li

Autonomous driving training requires a diverse range of datasets encompassing various traffic conditions, weather scenarios, and road types. Traditional data augmentation methods often struggle to generate datasets that represent rare…

计算机视觉与模式识别 · 计算机科学 2024-08-29 Yongjie Fu , Yunlong Li , Xuan Di

View-predictive generative models provide strong priors for lifting object-centric images and videos into 3D and 4D through rendering and score distillation objectives. A question then remains: what about lifting complete multi-object…

计算机视觉与模式识别 · 计算机科学 2024-05-24 Wen-Hsuan Chu , Lei Ke , Katerina Fragkiadaki

Despite remarkable achievements in video synthesis, achieving granular control over complex dynamics, such as nuanced movement among multiple interacting objects, still presents a significant hurdle for dynamic world modeling, compounded by…

计算机视觉与模式识别 · 计算机科学 2024-03-21 Pengxiang Li , Kai Chen , Zhili Liu , Ruiyuan Gao , Lanqing Hong , Guo Zhou , Hua Yao , Dit-Yan Yeung , Huchuan Lu , Xu Jia