English
Related papers

Related papers: MotionClone: Training-Free Motion Cloning for Cont…

200 papers

Human motion generation has been widely studied due to its crucial role in areas such as digital humans and humanoid robot control. However, many current motion generation approaches disregard physics constraints, frequently resulting in…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Zhuo Li , Mingshuang Luo , Ruibing Hou , Xin Zhao , Hao Liu , Hong Chang , Zimo Liu , Chen Li

We introduce a novel method for controlling a motion sequence using an arbitrary temporal control sequence using temporal alignment. Temporal alignment of motion has gained significant attention owing to its applications in motion control…

Graphics · Computer Science 2025-11-26 Naoki Agata , Takeo Igarashi

Existing video generation models struggle to follow complex text prompts and synthesize multiple objects, raising the need for additional grounding input for improved controllability. In this work, we propose to decompose videos into visual…

Computer Vision and Pattern Recognition · Computer Science 2025-01-15 Weixi Feng , Chao Liu , Sifei Liu , William Yang Wang , Arash Vahdat , Weili Nie

A natural approach to generative modeling of videos is to represent them as a composition of moving objects. Recent works model a set of 2D sprites over a slowly-varying background, but without considering the underlying 3D scene that gives…

Computer Vision and Pattern Recognition · Computer Science 2021-03-26 Paul Henderson , Christoph H. Lampert

Image animation brings life to the static object in the source image according to the driving video. Recent works attempt to perform motion transfer on arbitrary objects through unsupervised methods without using a priori knowledge.…

Computer Vision and Pattern Recognition · Computer Science 2022-03-30 Jian Zhao , Hui Zhang

Despite tremendous progress in dexterous manipulation, current visuomotor policies remain fundamentally limited by two challenges: they struggle to generalize under perceptual or behavioral distribution shifts, and their performance is…

Robotics · Computer Science 2025-08-04 Junbang Liang , Pavel Tokmakov , Ruoshi Liu , Sruthi Sudhakar , Paarth Shah , Rares Ambrus , Carl Vondrick

We propose a strong baseline model for unsupervised feature learning using video data. By learning to predict missing frames or extrapolate future frames from an input video sequence, the model discovers both spatial and temporal…

Machine Learning · Computer Science 2016-05-05 MarcAurelio Ranzato , Arthur Szlam , Joan Bruna , Michael Mathieu , Ronan Collobert , Sumit Chopra

Generating diverse and natural human motion is one of the long-standing goals for creating intelligent characters in the animated world. In this paper, we propose a self-supervised method for generating long-range, diverse and plausible…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Jingwei Xu , Huazhe Xu , Bingbing Ni , Xiaokang Yang , Xiaolong Wang , Trevor Darrell

Generating human motion with precise spatial control is a challenging problem. Existing approaches often require task-specific training or slow optimization, and enforcing hard constraints frequently disrupts motion naturalness. Building on…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Akihisa Watanabe , Qing Yu , Edgar Simo-Serra , Kent Fujiwara

We present FlexTraj, a framework for image-to-video generation with flexible point trajectory control. FlexTraj introduces a unified point-based motion representation that encodes each point with a segmentation ID, a temporally consistent…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Zhiyuan Zhang , Can Wang , Dongdong Chen , Jing Liao

Recent advancements in video generation have enabled the development of ``world models'' capable of simulating potential futures for robotics and planning. However, specifying precise goals for these models remains a challenge; text…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Nate Gillman , Yinghua Zhou , Zitian Tang , Evan Luo , Arjan Chakravarthy , Daksh Aggarwal , Michael Freeman , Charles Herrmann , Chen Sun

Human motions are compositional: complex behaviors can be described as combinations of simpler primitives. However, existing approaches primarily focus on forward modeling, e.g., learning holistic mappings from text to motion or composing a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Jianrong Zhang , Hehe Fan , Yi Yang

Video Diffusion Models have been developed for video generation, usually integrating text and image conditioning to enhance control over the generated content. Despite the progress, ensuring consistency across frames remains a challenge,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Tian Xia , Xuweiyi Chen , Sihan Xu

Generating motion sequences conforming to a target style while adhering to the given content prompts requires accommodating both the content and style. In existing methods, the information usually only flows from style to content, which may…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Zhe Li , Yisheng He , Lei Zhong , Weichao Shen , Qi Zuo , Lingteng Qiu , Zilong Dong , Laurence Tianruo Yang , Weihao Yuan

Text-to-video diffusion models are notoriously limited in their ability to model temporal aspects such as motion, physics, and dynamic interactions. Existing approaches address this limitation by retraining the model or introducing external…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Ariel Shaulov , Itay Hazan , Lior Wolf , Hila Chefer

Pretraining robust vision or multimodal foundation models (e.g., CLIP) relies on large-scale datasets that may be noisy, potentially misaligned, and have long-tail distributions. Previous works have shown promising results in augmenting…

Computer Vision and Pattern Recognition · Computer Science 2024-10-17 Qingqing Cao , Mahyar Najibi , Sachin Mehta

We aim to control a robot to physically behave in the real world following any high-level language command like "cartwheel" or "kick". Although human motion datasets exist, this task remains particularly challenging since generative models…

Robotics · Computer Science 2024-05-21 Shusheng Xu , Huaijie Wang , Jiaxuan Gao , Yutao Ouyang , Chao Yu , Yi Wu

The image-to-video (I2V) generation is conditioned on the static image, which has been enhanced recently by the motion intensity as an additional control signal. These motion-aware models are appealing to generate diverse motion patterns,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Shuwei Shi , Biao Gong , Xi Chen , Dandan Zheng , Shuai Tan , Zizheng Yang , Yuyuan Li , Jingwen He , Kecheng Zheng , Jingdong Chen , Ming Yang , Yinqiang Zheng

Recent works have successfully extended large-scale text-to-image models to the video domain, producing promising results but at a high computational cost and requiring a large amount of video data. In this work, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Bo Peng , Xinyuan Chen , Yaohui Wang , Chaochao Lu , Yu Qiao

Recent advances in text-to-video diffusion models have enabled high-fidelity and temporally coherent videos synthesis. However, current models are predominantly optimized for single-event generation. When handling multi-event prompts,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Qianxun Xu , Chenxi Song , Yujun Cai , Chi Zhang