English
Related papers

Related papers: Motion-Conditioned Image Animation for Video Editi…

200 papers

Motion is a salient cue to recognize actions in video. Modern action recognition models leverage motion information either explicitly by using optical flow as input or implicitly by means of 3D convolutional filters that simultaneously…

Computer Vision and Pattern Recognition · Computer Science 2020-05-28 Heng Wang , Du Tran , Lorenzo Torresani , Matt Feiszli

Text-to-image diffusion models have demonstrated remarkable progress in synthesizing high-quality images from text prompts, which boosts researches on prompt-based image editing that edits a source image according to a target prompt.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Kejie Wang , Xuemeng Song , Meng Liu , Jin Yuan , Weili Guan

Text-guided motion editing enables high-level semantic control and iterative modifications beyond traditional keyframe animation. Existing methods rely on limited pre-collected training triplets, which severely hinders their versatility in…

Computer Vision and Pattern Recognition · Computer Science 2025-09-15 Nan Jiang , Hongjie Li , Ziye Yuan , Zimo He , Yixin Chen , Tengyu Liu , Yixin Zhu , Siyuan Huang

Driven by the upsurge progress in text-to-image (T2I) generation models, text-to-video (T2V) generation has experienced a significant advance as well. Accordingly, tasks such as modifying the object or changing the style in a video have…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Yeji Song , Wonsik Shin , Junsoo Lee , Jeesoo Kim , Nojun Kwak

State-of-the-art video action classifiers often suffer from overfitting. They tend to be biased towards specific objects and scene cues, rather than the foreground action content, leading to sub-optimal generalization performances. Recent…

Computer Vision and Pattern Recognition · Computer Science 2020-12-08 Sangdoo Yun , Seong Joon Oh , Byeongho Heo , Dongyoon Han , Jinhyung Kim

In this work we propose to utilize information about human actions to improve pose estimation in monocular videos. To this end, we present a pictorial structure model that exploits high-level information about activities to incorporate…

Computer Vision and Pattern Recognition · Computer Science 2017-02-13 Umar Iqbal , Martin Garbade , Juergen Gall

In this paper, we propose the LoRA of Change (LoC) framework for image editing with visual instructions, i.e., before-after image pairs. Compared to the ambiguities, insufficient specificity, and diverse interpretations of natural language,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Xue Song , Jiequan Cui , Hanwang Zhang , Jiaxin Shi , Jingjing Chen , Chi Zhang , Yu-Gang Jiang

Motion correction aims to prevent motion artefacts which may be caused by respiration, heartbeat, or head movements for example. In a preliminary step, the measured data is divided in gates corresponding to motion states, and displacement…

Optimization and Control · Mathematics 2024-10-15 Claire Delplancke , Kris Thielemans , Matthias J. Ehrhardt

Recent vision transformer based video models mostly follow the ``image pre-training then finetuning" paradigm and have achieved great success on multiple video benchmarks. However, full finetuning such a video model could be computationally…

Computer Vision and Pattern Recognition · Computer Science 2023-02-07 Taojiannan Yang , Yi Zhu , Yusheng Xie , Aston Zhang , Chen Chen , Mu Li

Current diffusion-based face animation methods generally adopt a ReferenceNet (a copy of U-Net) and a large amount of curated self-acquired data to learn appearance features, as robust appearance features are vital for ensuring temporal…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Yue Han , Junwei Zhu , Yuxiang Feng , Xiaozhong Ji , Keke He , Xiangtai Li , zhucun xue , Yong Liu

Recent advances in image editing, driven by image diffusion models, have shown remarkable progress. However, significant challenges remain, as these models often struggle to follow complex edit instructions accurately and frequently…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Noam Rotstein , Gal Yona , Daniel Silver , Roy Velich , David Bensaïd , Ron Kimmel

In this work, we propose a novel and efficient method for articulated human pose estimation in videos using a convolutional network architecture, which incorporates both color and motion features. We propose a new human body pose dataset,…

Computer Vision and Pattern Recognition · Computer Science 2014-09-30 Arjun Jain , Jonathan Tompson , Yann LeCun , Christoph Bregler

Understanding instructional videos requires recognizing fine-grained actions and modeling their temporal relations, which remains challenging for current Video Foundation Models (VFMs). This difficulty stems from noisy web supervision and a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Zhuoyi Yang , Jiapeng Yu , Reuben Tan , Boyang Li , Huijuan Xu

Image generation and editing have seen a great deal of advancements with the rise of large-scale diffusion models that allow user control of different modalities such as text, mask, depth maps, etc. However, controlled editing of videos…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 AmirHossein Zamani , Amir G. Aghdam , Tiberiu Popa , Eugene Belilovsky

Text-based video editing has recently attracted considerable interest in changing the style or replacing the objects with a similar structure. Beyond this, we demonstrate that properties such as shape, size, location, motion, etc., can also…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Yue Ma , Xiaodong Cun , Sen Liang , Jinbo Xing , Yingqing He , Chenyang Qi , Siran Chen , Qifeng Chen

Advancements in attention mechanisms have led to significant performance improvements in a variety of areas in machine learning due to its ability to enable the dynamic modeling of temporal sequences. A particular area in computer vision…

Computer Vision and Pattern Recognition · Computer Science 2021-12-14 Brennan Gebotys , Alexander Wong , David A. Clausi

Video editing is a challenging task that requires manipulating videos on both the spatial and temporal dimensions. Existing methods for video editing mainly focus on changing the appearance or style of the objects in the video, while…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Yao Teng , Enze Xie , Yue Wu , Haoyu Han , Zhenguo Li , Xihui Liu

A unified video and action model holds significant promise for robotics, where videos provide rich scene information for action prediction, and actions provide dynamics information for video prediction. However, effectively combining video…

Robotics · Computer Science 2025-04-28 Shuang Li , Yihuai Gao , Dorsa Sadigh , Shuran Song

Recent studies on motion estimation have advocated an optimized motion representation that is globally consistent across the entire video, preferably for every pixel. This is challenging as a uniform representation may not account for the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Rui Li , Dong Liu

Image animation has seen significant progress, driven by the powerful generative capabilities of diffusion models. However, maintaining appearance consistency with static input images and mitigating abrupt motion transitions in generated…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Xin Ma , Yaohui Wang , Genyun Jia , Xinyuan Chen , Tien-Tsin Wong , Cunjian Chen