中文
相关论文

相关论文: MoCA-Video: Motion-Aware Concept Alignment for Con…

200 篇论文

The goal of this paper is to optimize the training process of diffusion-based text-to-speech models. While recent studies have achieved remarkable advancements, their training demands substantial time and computational costs, largely due to…

音频与语音处理 · 电气工程与系统科学 2025-06-02 Jeongsoo Choi , Zhikang Niu , Ji-Hoon Kim , Chunhui Wang , Joon Son Chung , Xie Chen

Large-scale pre-trained diffusion models have exhibited remarkable capabilities in diverse video generations. Given a set of video clips of the same motion concept, the task of Motion Customization is to adapt existing text-to-video…

计算机视觉与模式识别 · 计算机科学 2023-10-13 Rui Zhao , Yuchao Gu , Jay Zhangjie Wu , David Junhao Zhang , Jiawei Liu , Weijia Wu , Jussi Keppo , Mike Zheng Shou

Recent text-to-video diffusion models have achieved impressive progress. In practice, users often desire the ability to control object motion and camera movement independently for customized video creation. However, current methods lack the…

计算机视觉与模式识别 · 计算机科学 2024-05-07 Shiyuan Yang , Liang Hou , Haibin Huang , Chongyang Ma , Pengfei Wan , Di Zhang , Xiaodong Chen , Jing Liao

Long video generation involves generating extended videos using models trained on short videos, suffering from distribution shifts due to varying frame counts. It necessitates the use of local information from the original short frames to…

计算机视觉与模式识别 · 计算机科学 2025-05-05 Jiangtong Tan , Hu Yu , Jie Huang , Jie Xiao , Feng Zhao

In recent years, video semantic segmentation has made great progress with advanced deep neural networks. However, there still exist two main challenges \ie, information inconsistency and computation cost. To deal with the two difficulties,…

计算机视觉与模式识别 · 计算机科学 2023-04-19 Jinming Su , Ruihong Yin , Shuaibin Zhang , Junfeng Luo

Temporal modeling and spatio-temporal collaboration are pivotal techniques for video-based human pose estimation. Most state-of-the-art methods adopt optical flow or temporal difference, learning local visual content correspondence across…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Runyang Feng , Haoming Chen

Without incurring significant computational overhead, train-free long video generation aims to enable foundation video generation models to produce longer videos. Frame-level autoregressive frameworks, e.g., FIFO-diffusion, offer the…

计算机视觉与模式识别 · 计算机科学 2026-05-19 X. Feng , J. Zhu , M. Wu , C. Chen , F. Mao , H. Guo , J. Wu , X. Chu , K. Huang

Video diffusion models are moving beyond short, plausible clips toward world simulators that must remain consistent under camera motion, revisits, and intervention. Yet spatial memory remains a key bottleneck: explicit 3D structures can…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Wei Yu , Runjia Qian , Yumeng Li , Liquan Wang , Songheng Yin , Sri Siddarth Chakaravarthy P , Dennis Anthony , Yang Ye , Yidi Li , Weiwei Wan , Animesh Garg

Recent advancements in video generation have been remarkable, yet many existing methods struggle with issues of consistency and poor text-video alignment. Moreover, the field lacks effective techniques for text-guided video inpainting, a…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Bojia Zi , Shihao Zhao , Xianbiao Qi , Jianan Wang , Yukai Shi , Qianyu Chen , Bin Liang , Kam-Fai Wong , Lei Zhang

Diffusion-based text-to-image (T2I) models have demonstrated remarkable results in global video editing tasks. However, their focus is primarily on global video modifications, and achieving desired attribute-specific changes remains a…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Haoyu Zheng , Wenqiao Zhang , Zheqi Lv , Yu Zhong , Yang Dai , Jianxiang An , Yongliang Shen , Juncheng Li , Dongping Zhang , Siliang Tang , Yueting Zhuang

We address the task of multi-view image editing from sparse input views, where the inputs can be seen as a mix of images capturing the scene from different viewpoints. The goal is to modify the scene according to a textual instruction while…

计算机视觉与模式识别 · 计算机科学 2025-11-20 Daniel Gilo , Or Litany

We propose a method for adding sound-guided visual effects to specific regions of videos with a zero-shot setting. Animating the appearance of the visual effect is challenging because each frame of the edited video should have visual…

计算机视觉与模式识别 · 计算机科学 2023-04-17 Seung Hyun Lee , Sieun Kim , Innfarn Yoo , Feng Yang , Donghyeon Cho , Youngseo Kim , Huiwen Chang , Jinkyu Kim , Sangpil Kim

Multimodal video summarization requires visual features that align semantically with language generation. Traditional approaches rely on CNN features trained for object classification, which represent visual concepts as discrete categories…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Maham Nazir , Muhammad Aqeel , Richong Zhang , Francesco Setti

Large-scale text-to-image (T2I) diffusion models have been extended for text-guided video editing, yielding impressive zero-shot video editing performance. Nonetheless, the generated videos usually show spatial irregularities and temporal…

计算机视觉与模式识别 · 计算机科学 2024-08-29 Yuanzhi Wang , Yong Li , Xiaoya Zhang , Xin Liu , Anbo Dai , Antoni B. Chan , Zhen Cui

Reasoning Video Object Segmentation (ReasonVOS) is a challenging task that requires stable object segmentation across video sequences using implicit and complex textual inputs. Previous methods fine-tune Multimodal Large Language Models…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Zhengtong Zhu , Jiaqing Fan , Zhixuan Liu , Fanzhang Li

Typical deep neural video compression networks usually follow the hybrid approach of classical video coding that contains two separate modules: motion coding and residual coding. In addition, a symmetric auto-encoder is often used as a…

图像与视频处理 · 电气工程与系统科学 2024-11-27 Van Thang Nguyen

Large-scale text-to-image diffusion models achieve unprecedented success in image generation and editing. However, how to extend such success to video editing is unclear. Recent initial attempts at video editing require significant…

计算机视觉与模式识别 · 计算机科学 2024-01-05 Wen Wang , Yan Jiang , Kangyang Xie , Zide Liu , Hao Chen , Yue Cao , Xinlong Wang , Chunhua Shen

Recently, several works tackled the video editing task fostered by the success of large-scale text-to-image generative models. However, most of these methods holistically edit the frame using the text, exploiting the prior given by…

计算机视觉与模式识别 · 计算机科学 2024-01-08 Elia Peruzzo , Vidit Goel , Dejia Xu , Xingqian Xu , Yifan Jiang , Zhangyang Wang , Humphrey Shi , Nicu Sebe

Video demoireing aims to remove undesirable interference patterns that arise during the capture of screen content, restoring artifact-free frames while maintaining temporal consistency. Existing video demoireing methods typically utilize…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Shuning Xu , Xina Liu , Binbin Song , Xiangyu Chen , Qiubo Chen , Jiantao Zhou

StyleMamba has recently demonstrated efficient text-driven image style transfer by leveraging state-space models (SSMs) and masked directional losses. In this paper, we extend the StyleMamba framework to handle video sequences. We propose…

图形学 · 计算机科学 2025-07-31 Chao Li , Minsu Park , Cristina Rossi , Zhuang Li