中文
相关论文

相关论文: I2VControl: Disentangled and Unified Video Motion …

200 篇论文

Temporal modeling on regular respiration-induced motions is crucial to image-guided clinical applications. Existing methods cannot simulate temporal motions unless high-dose imaging scans including starting and ending frames exist…

计算机视觉与模式识别 · 计算机科学 2025-06-26 Xin You , Minghui Zhang , Hanxiao Zhang , Jie Yang , Nassir Navab

As virtual reality gains popularity, the demand for controllable creation of immersive and dynamic omnidirectional videos (ODVs) is increasing. While previous text-to-ODV generation methods achieve impressive results, they struggle with…

计算机视觉与模式识别 · 计算机科学 2024-12-13 Weiqi Li , Shijie Zhao , Chong Mou , Xuhan Sheng , Zhenyu Zhang , Qian Wang , Junlin Li , Li Zhang , Jian Zhang

Controllable speech synthesis aims to control the style of generated speech using reference input, which can be of various modalities. Existing face-based methods struggle with robustness and generalization due to data quality constraints,…

声音 · 计算机科学 2025-06-27 Rui Niu , Weihao Wu , Jie Chen , Long Ma , Zhiyong Wu

We consider the task of Image-to-Video (I2V) generation, which involves transforming static images into realistic video sequences based on a textual description. While recent advancements produce photorealistic outputs, they frequently…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Guy Yariv , Yuval Kirstain , Amit Zohar , Shelly Sheynin , Yaniv Taigman , Yossi Adi , Sagie Benaim , Adam Polyak

While recent advancements in text-to-video diffusion models enable high-quality short video generation from a single prompt, generating real-world long videos in a single pass remains challenging due to limited data and high computational…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Subin Kim , Seoung Wug Oh , Jui-Hsien Wang , Joon-Young Lee , Jinwoo Shin

Diffusion models have demonstrated great success in text-to-video (T2V) generation. However, existing methods may face challenges when handling complex (long) video generation scenarios that involve multiple objects or dynamic changes in…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Ye Tian , Ling Yang , Haotian Yang , Yuan Gao , Yufan Deng , Jingmin Chen , Xintao Wang , Zhaochen Yu , Xin Tao , Pengfei Wan , Di Zhang , Bin Cui

Videos of actions are complex signals containing rich compositional structure in space and time. Current video generation methods lack the ability to condition the generation on multiple coordinated and potentially simultaneous timed…

计算机视觉与模式识别 · 计算机科学 2021-06-14 Amir Bar , Roei Herzig , Xiaolong Wang , Anna Rohrbach , Gal Chechik , Trevor Darrell , Amir Globerson

Within recent approaches to text-to-video (T2V) generation, achieving controllability in the synthesized video is often a challenge. Typically, this issue is addressed by providing low-level per-frame guidance in the form of edge maps,…

计算机视觉与模式识别 · 计算机科学 2024-04-10 Wan-Duo Kurt Ma , J. P. Lewis , W. Bastiaan Kleijn

Identity-preserving text-to-video (IPT2V) generation, which aims to create high-fidelity videos with consistent human identity, has become crucial for downstream applications. However, current end-to-end frameworks suffer a critical…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Yuji Wang , Moran Li , Xiaobin Hu , Ran Yi , Jiangning Zhang , Han Feng , Weijian Cao , Yabiao Wang , Chengjie Wang , Lizhuang Ma

Recent advances in diffusion models have showcased promising results in the text-to-video (T2V) synthesis task. However, as these T2V models solely employ text as the guidance, they tend to struggle in modeling detailed temporal dynamics.…

计算机视觉与模式识别 · 计算机科学 2023-05-24 Seungwoo Lee , Chaerin Kong , Donghyeon Jeon , Nojun Kwak

Effective and generalizable control in video generation remains a significant challenge. While many methods rely on ambiguous or task-specific signals, we argue that a fundamental disentanglement of "appearance" and "motion" provides a more…

计算机视觉与模式识别 · 计算机科学 2026-02-16 Mingzhi Sheng , Zekai Gu , Peng Li , Cheng Lin , Hao-Xiang Guo , Ying-Cong Chen , Yuan Liu

Human-centric motion control in video generation remains a critical challenge, particularly when jointly controlling camera movements and human poses in scenarios like the iconic Grammy Glambot moment. While recent video diffusion models…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Ruineng Li , Daitao Xing , Huiming Sun , Yuanzhou Ha , Jinglin Shen , Chiuman Ho

Diffusion-based \textit{image-to-video} (I2V) generation has become a central direction in generative models by turning a reference image, with optional conditions, into a temporally coherent video. Compared with broader video generation…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Xianlong Wang , Wenbo Pan , Shijia Zhou , Ke Li , Yuqi Wang , Zeyu Ye , Hangtao Zhang , Leo Yu Zhang , Xiaohua Jia

Using synthesized images to boost the performance of perception models is a long-standing research challenge in computer vision. It becomes more eminent in visual-centric autonomous driving systems with multi-view cameras as some long-tail…

计算机视觉与模式识别 · 计算机科学 2023-09-26 Kairui Yang , Enhui Ma , Jibin Peng , Qing Guo , Di Lin , Kaicheng Yu

Text-image-to-video (TI2V) generation is a critical problem for controllable video generation using both semantic and visual conditions. Most existing methods typically add visual conditions to text-to-video (T2V) foundation models by…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Bolin Lai , Sangmin Lee , Xu Cao , Xiang Li , James M. Rehg

Recent advancements in personalized Text-to-Video (T2V) generation have made significant strides in synthesizing character-specific content. However, these methods face a critical limitation: the inability to perform fine-grained control…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Haopeng Fang , Di Qiu , Binjie Mao , He Tang

Controlling video and audio generation requires diverse modalities, from depth and pose to camera trajectories and audio transformations, yet existing approaches either train a single monolithic model for a fixed set of controls or…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Matan Ben-Yosef , Tavi Halperin , Naomi Ken Korem , Mohammad Salama , Harel Cain , Asaf Joseph , Anthony Chen , Urska Jelercic , Ofir Bibi

The goal of conditional image-to-video (cI2V) generation is to create a believable new video by beginning with the condition, i.e., one image and text.The previous cI2V generation methods conventionally perform in RGB pixel space, with…

计算机视觉与模式识别 · 计算机科学 2023-12-15 Cuifeng Shen , Yulu Gan , Chen Chen , Xiongwei Zhu , Lele Cheng , Tingting Gao , Jinzhi Wang

We introduce InstructVid2Vid, an end-to-end diffusion-based methodology for video editing guided by human language instructions. Our approach empowers video manipulation guided by natural language directives, eliminating the need for…

计算机视觉与模式识别 · 计算机科学 2024-05-30 Bosheng Qin , Juncheng Li , Siliang Tang , Tat-Seng Chua , Yueting Zhuang

We introduce a framework that enables both multi-view character consistency and 3D camera control in video diffusion models through a novel customization data pipeline. We train the character consistency component with recorded volumetric…