English
Related papers

Related papers: UniMLVG: Unified Framework for Multi-view Long Vid…

200 papers

World models have demonstrated significant promise for data synthesis in autonomous driving. However, existing methods predominantly concentrate on single-modality generation, typically focusing on either multi-camera video or LiDAR…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Guosheng Zhao , Yaozeng Wang , Xiaofeng Wang , Zheng Zhu , Tingdong Yu , Guan Huang , Yongchen Zai , Ji Jiao , Changliang Xue , Xiaole Wang , Zhen Yang , Futang Zhu , Xingang Wang

We propose UniMo, an innovative autoregressive model for joint modeling of 2D human videos and 3D human motions within a unified framework, enabling simultaneous generation and understanding of these two modalities for the first time.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Youxin Pang , Yong Zhang , Ruizhi Shao , Xiang Deng , Feng Gao , Xu Xiaoming , Xiaoming Wei , Yebin Liu

The emergence of Diffusion Transformers (DiT) has brought significant advancements to video generation, especially in text-to-video and image-to-video tasks. Although video generation is widely applied in various fields, most existing…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Sen Liang , Zhentao Yu , Zhengguang Zhou , Teng Hu , Hongmei Wang , Yi Chen , Qin Lin , Yuan Zhou , Xin Li , Qinglin Lu , Zhibo Chen

Video foundation models aim to integrate video understanding, generation, editing, and instruction following within a single framework, making them a central direction for next-generation multimodal systems. However, existing evaluation…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Jianhui Wei , Xiaotian Zhang , Yichen Li , Yuan Wang , Yan Zhang , Ziyi Chen , Zhihang Tang , Wei Xu , Zuozhu Liu

In recent years there have been remarkable breakthroughs in image-to-video generation. However, the 3D consistency and camera controllability of generated frames have remained unsolved. Recent studies have attempted to incorporate camera…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Dejia Xu , Yifan Jiang , Chen Huang , Liangchen Song , Thorsten Gernoth , Liangliang Cao , Zhangyang Wang , Hao Tang

Recent successful video generation systems that predict and create realistic automotive driving scenes from short video inputs assign tokenization, future state prediction (world model), and video decoding to dedicated models. These…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Björn Möller , Zhengyang Li , Malte Stelzer , Thomas Graave , Fabian Bettels , Muaaz Ataya , Tim Fingscheidt

The real world is dynamic, yet most image fusion methods process static frames independently, ignoring temporal correlations in videos and leading to flickering and temporal inconsistency. To address this, we propose Unified Video Fusion…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Zixiang Zhao , Haowen Bai , Bingxin Ke , Yukun Cui , Lilun Deng , Yulun Zhang , Kai Zhang , Konrad Schindler

In autonomous driving, predicting future events in advance and evaluating the foreseeable risks empowers autonomous vehicles to better plan their actions, enhancing safety and efficiency on the road. To this end, we propose Drive-WM, the…

Computer Vision and Pattern Recognition · Computer Science 2023-11-30 Yuqi Wang , Jiawei He , Lue Fan , Hongxin Li , Yuntao Chen , Zhaoxiang Zhang

With recent advances in Multimodal Large Language Models (MLLMs) showing strong visual understanding and reasoning, interest is growing in using them to improve the editing performance of diffusion models. Despite rapid progress, most…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Chong Mou , Qichao Sun , Yanze Wu , Pengze Zhang , Xinghui Li , Fulong Ye , Songtao Zhao , Qian He

Video generation has witnessed remarkable progress with the advent of deep generative models, particularly diffusion models. While existing methods excel in generating high-quality videos from text prompts or single images, personalized…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Yufan Deng , Xun Guo , Yizhi Wang , Jacob Zhiyuan Fang , Angtian Wang , Shenghai Yuan , Yiding Yang , Bo Liu , Haibin Huang , Chongyang Ma

The field of advanced text-to-image generation is witnessing the emergence of unified frameworks that integrate powerful text encoders, such as CLIP and T5, with Diffusion Transformer backbones. Although there have been efforts to control…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Liang Chen , Shuai Bai , Wenhao Chai , Weichu Xie , Haozhe Zhao , Leon Vinci , Junyang Lin , Baobao Chang

In recent years, artificial intelligence (AI)-driven video generation has gained significant attention. Consequently, there is a growing need for accurate video quality assessment (VQA) metrics to evaluate the perceptual quality of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Zhichao Zhang , Wei Sun , Xinyue Li , Jun Jia , Xiongkuo Min , Zicheng Zhang , Chunyi Li , Zijian Chen , Puyi Wang , Fengyu Sun , Shangling Jui , Guangtao Zhai

While diffusion models have shown impressive performance in 2D image/video generation, diffusion-based Text-to-Multi-view-Video (T2MVid) generation remains underexplored. The new challenges posed by T2MVid generation lie in the lack of…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Bing Li , Cheng Zheng , Wenxuan Zhu , Jinjie Mai , Biao Zhang , Peter Wonka , Bernard Ghanem

Text-to-Image (T2I) diffusion models have shown impressive results in generating visually compelling images following user prompts. Building on this, various methods further fine-tune the pre-trained T2I model for specific tasks. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-24 Tsu-Jui Fu , Yusu Qian , Chen Chen , Wenze Hu , Zhe Gan , Yinfei Yang

While text-to-video diffusion models have advanced significantly, creating coherent long-form content remains unreliable due to stochastic sampling artifacts. This necessitates generating multiple candidates, yet verifying them creates a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Daewon Yoon , Hyeongseok Lee , Wonsik Shin , Sangyu Han , Nojun Kwak

We present UniFuture, a unified 4D Driving World Model designed to simulate the dynamic evolution of the 3D physical world. Unlike existing driving world models that focus solely on 2D pixel-level video generation (lacking geometry) or…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Dingkang Liang , Dingyuan Zhang , Xin Zhou , Sifan Tu , Tianrui Feng , Xiaofan Li , Yumeng Zhang , Mingyang Du , Xiao Tan , Xiang Bai

With the rapid growth of video generative models (VGMs), it is essential to develop reliable and comprehensive automatic metrics for AI-generated videos (AIGVs). Existing methods either use off-the-shelf models optimized for other tasks or…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Yuanxin Liu , Rui Zhu , Shuhuai Ren , Jiacong Wang , Haoyuan Guo , Xu Sun , Lu Jiang

Modeling scenes using video generation models has garnered growing research interest in recent years. However, most existing approaches rely on perspective video models that synthesize only limited observations of a scene, leading to issues…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Yuheng Liu , Xin Lin , Xinke Li , Baihan Yang , Chen Wang , Kalyan Sunkavalli , Yannick Hold-Geoffroy , Hao Tan , Kai Zhang , Xiaohui Xie , Zifan Shi , Yiwei Hu

In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff , aiming to synthesize and comprehend multiple video visual content in a single diffusion model. To achieve this, OmniVDiff treats all video visual…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Dianbing Xi , Jiepeng Wang , Yuanzhi Liang , Xi Qiu , Yuchi Huo , Rui Wang , Chi Zhang , Xuelong Li

Generating multi-camera street-view videos is critical for augmenting autonomous driving datasets, addressing the urgent demand for extensive and varied data. Due to the limitations in diversity and challenges in handling lighting…

Computer Vision and Pattern Recognition · Computer Science 2024-08-09 Jiachen Lu , Ze Huang , Zeyu Yang , Jiahui Zhang , Li Zhang