中文
相关论文

相关论文: MAD: Motion Appearance Decoupling for efficient Dr…

200 篇论文

Motion customization aims to adapt the diffusion model (DM) to generate videos with the motion specified by a set of video clips with the same motion concept. To realize this goal, the adaptation of DM should be possible to model the…

计算机视觉与模式识别 · 计算机科学 2025-09-09 Huijie Liu , Jingyun Wang , Shuai Ma , Jie Hu , Xiaoming Wei , Guoliang Kang

Existing video generation models predominantly emphasize appearance fidelity while exhibiting limited ability to synthesize complex human motions, such as whole-body movements, long-range dynamics, and fine-grained human-environment…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Haoyu Wang , Hao Tang , Donglin Di , Zhilu Zhang , Wangmeng Zuo , Feng Gao , Siwei Ma , Shiliang Zhang

Autonomous driving requires robust perception models trained on high-quality, large-scale multi-view driving videos for tasks like 3D object detection, segmentation and trajectory prediction. While world models provide a cost-effective…

计算机视觉与模式识别 · 计算机科学 2024-12-16 Zhuoran Yang , Xi Guo , Chenjing Ding , Chiyu Wang , Wei Wu

Image-to-Video (I2V) generation aims to synthesize a video clip according to a given image and condition (e.g., text). The key challenge of this task lies in simultaneously generating natural motions while preserving the original appearance…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Jie Tian , Xiaoye Qu , Zhenyi Lu , Wei Wei , Sichen Liu , Yu Cheng

A free-viewpoint, editable, and high-fidelity driving simulator is crucial for training and evaluating end-to-end autonomous driving systems. In this paper, we present GA-Drive, a novel simulation framework capable of generating camera…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Hao Zhang , Lue Fan , Qitai Wang , Wenbo Li , Zehuan Wu , Lewei Lu , Zhaoxiang Zhang , Hongsheng Li

Text-to-video diffusion models have enabled high-quality video synthesis, yet often fail to generate temporally coherent and physically plausible motion. A key reason is the models' insufficient understanding of complex motions that natural…

计算机视觉与模式识别 · 计算机科学 2025-10-23 Aritra Bhowmik , Denis Korzhenkov , Cees G. M. Snoek , Amirhossein Habibian , Mohsen Ghafoorian

World models, which predict future transitions from past observation and action sequences, have shown great promise for improving data efficiency in sequential decision-making. However, existing world models often require extensive…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Siqiao Huang , Jialong Wu , Qixing Zhou , Shangchen Miao , Mingsheng Long

Large-scale pre-trained diffusion models have exhibited remarkable capabilities in diverse video generations. Given a set of video clips of the same motion concept, the task of Motion Customization is to adapt existing text-to-video…

计算机视觉与模式识别 · 计算机科学 2023-10-13 Rui Zhao , Yuchao Gu , Jay Zhangjie Wu , David Junhao Zhang , Jiawei Liu , Weijia Wu , Jussi Keppo , Mike Zheng Shou

The video generation field has witnessed rapid improvements with the introduction of recent diffusion models. While these models have successfully enhanced appearance quality, they still face challenges in generating coherent and natural…

计算机视觉与模式识别 · 计算机科学 2025-04-21 Yaosi Hu , Zhenzhong Chen , Chong Luo

World models can foresee the outcomes of different actions, which is of paramount importance for autonomous driving. Nevertheless, existing driving world models still have limitations in generalization to unseen environments, prediction…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Shenyuan Gao , Jiazhi Yang , Li Chen , Kashyap Chitta , Yihang Qiu , Andreas Geiger , Jun Zhang , Hongyang Li

We propose a generative model that, given a coarsely edited image, synthesizes a photorealistic output that follows the prescribed layout. Our method transfers fine details from the original image and preserve the identity of its parts.…

计算机视觉与模式识别 · 计算机科学 2025-08-08 Hadi Alzayer , Zhihao Xia , Xuaner Zhang , Eli Shechtman , Jia-Bin Huang , Michael Gharbi

Generating human videos with realistic and controllable motions is a challenging task. While existing methods can generate visually compelling videos, they lack separate control over four key video elements: foreground subject, background…

计算机视觉与模式识别 · 计算机科学 2025-08-13 Jingyun Liang , Jingkai Zhou , Shikai Li , Chenjie Cao , Lei Sun , Yichen Qian , Weihua Chen , Fan Wang

Data-driven learning has advanced autonomous driving, yet task-specific models struggle with out-of-distribution scenarios due to their narrow optimization objectives and reliance on costly annotated data. We present DriveX, a…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Chen Shi , Shaoshuai Shi , Kehua Sheng , Bo Zhang , Li Jiang

World models that forecast environmental changes from actions are vital for autonomous driving models with strong generalization. The prevailing driving world model mainly build on video prediction model. Although these models can produce…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Jingcheng Ni , Yuxin Guo , Yichen Liu , Rui Chen , Lewei Lu , Zehuan Wu

Large-scale generative models have achieved remarkable success in a number of domains. However, for sequential decision-making problems, such as robotics, action-labelled data is often scarce and therefore scaling-up foundation models for…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Marc Rigter , Tarun Gupta , Agrin Hilmkil , Chao Ma

Recent successes in autoregressive (AR) generation models, such as the GPT series in natural language processing, have motivated efforts to replicate this success in visual tasks. Some works attempt to extend this approach to autonomous…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Xiaotao Hu , Wei Yin , Mingkai Jia , Junyuan Deng , Xiaoyang Guo , Qian Zhang , Xiaoxiao Long , Ping Tan

The field of autonomous driving is experiencing a surge of interest in world models, which aim to predict potential future scenarios based on historical observations. In this paper, we introduce DFIT-OccWorld, an efficient 3D occupancy…

计算机视觉与模式识别 · 计算机科学 2024-12-19 Haiming Zhang , Ying Xue , Xu Yan , Jiacheng Zhang , Weichao Qiu , Dongfeng Bai , Bingbing Liu , Shuguang Cui , Zhen Li

Latent Diffusion Models (LDMs) enable high-quality image synthesis while avoiding excessive compute demands by training a diffusion model in a compressed lower-dimensional latent space. Here, we apply the LDM paradigm to high-resolution…

计算机视觉与模式识别 · 计算机科学 2023-12-29 Andreas Blattmann , Robin Rombach , Huan Ling , Tim Dockhorn , Seung Wook Kim , Sanja Fidler , Karsten Kreis

World models have gained significant attention as a promising approach for autonomous driving. By emulating human-like perception and decision-making processes, these models can predict and adapt to dynamic environments. Existing methods…

机器人学 · 计算机科学 2025-12-03 Huiqian Li , Wei Pan , Haodong Zhang , Jin Huang , Zhihua Zhong

Dynamic Mode Decomposition (DMD) is a numerical method that seeks to fit timeseries data to a linear dynamical system. In doing so, DMD decomposes dynamic data into spatially coherent modes that evolve in time according to exponential…

计算机视觉与模式识别 · 计算机科学 2026-02-26 Marco Mignacca , Simone Brugiapaglia , Jason J. Bramburger
‹ 上一页 1 2 3 10 下一页 ›