English
Related papers

Related papers: EchoMotion: Unified Human Video and Motion Generat…

200 papers

Learning directly from human demonstration videos is a key milestone toward scalable and generalizable robot learning. Yet existing methods rely on intermediate representations such as keypoints or trajectories, introducing information loss…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Yiren Song , Cheng Liu , Weijia Mao , Mike Zheng Shou

Several video-based 3D pose and shape estimation algorithms have been proposed to resolve the temporal inconsistency of single-image-based methods. However it still remains challenging to have stable and accurate reconstruction. In this…

Computer Vision and Pattern Recognition · Computer Science 2021-10-25 Ziwen Li , Bo Xu , Han Huang , Cheng Lu , Yandong Guo

Human pose, action, and motion generation are critical for applications in digital humans, character animation, and humanoid robotics. However, many existing methods struggle to produce physically plausible movements that are consistent…

Computer Vision and Pattern Recognition · Computer Science 2025-07-16 Zixi Kang , Xinghan Wang , Yadong Mu

The field has made significant progress in synthesizing realistic human motion driven by various modalities. Yet, the need for different methods to animate various body parts according to different control signals limits the scalability of…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Zixiang Zhou , Yu Wan , Baoyuan Wang

Generative modeling of human motion has broad applications in computer animation, virtual reality, and robotics. Conventional approaches develop separate models for different motion synthesis tasks, and typically use a model of a small size…

Computer Vision and Pattern Recognition · Computer Science 2022-12-07 Jianxin Ma , Shuai Bai , Chang Zhou

We present a unified perspective on tackling various human-centric video tasks by learning human motion representations from large-scale and heterogeneous data resources. Specifically, we propose a pretraining stage in which a motion…

Computer Vision and Pattern Recognition · Computer Science 2023-08-15 Wentao Zhu , Xiaoxuan Ma , Zhaoyang Liu , Libin Liu , Wayne Wu , Yizhou Wang

Achieving both behavioral similarity and appropriateness in human-like motion generation for humanoid robot remains an open challenge, further compounded by the lack of cross-embodiment adaptability. To address this problem, we propose…

Robotics · Computer Science 2025-08-27 Shipeng Lyu , Fangyuan Wang , Weiwei Lin , Luhao Zhu , David Navarro-Alarcon , Guodong Guo

Video Frame Interpolation (VFI) aims to synthesize intermediate frames between existing frames to enhance visual smoothness and quality. Beyond the conventional methods based on the reconstruction loss, recent works have employed generative…

Computer Vision and Pattern Recognition · Computer Science 2024-12-20 Jaihyun Lew , Jooyoung Choi , Chaehun Shin , Dahuin Jung , Sungroh Yoon

The visual world offers a critical axis for advancing foundation models beyond language. Despite growing interest in this direction, the design space for native multimodal models remains opaque. We provide empirical clarity through…

Text-driven human motion generation is an emerging task in animation and humanoid robot design. Existing algorithms directly generate the full sequence which is computationally expensive and prone to errors as it does not pay special…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Zichen Geng , Caren Han , Zeeshan Hayder , Jian Liu , Mubarak Shah , Ajmal Mian

Autoregressive large language models (LLMs) have unified a vast range of language tasks, inspiring preliminary efforts in autoregressive (AR) video generation. Existing AR video generators either diverge from standard LLM architectures,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Hangjie Yuan , Weihua Chen , Jun Cen , Hu Yu , Jingyun Liang , Shuning Chang , Zhihui Lin , Tao Feng , Pengwei Liu , Jiazheng Xing , Hao Luo , Jiasheng Tang , Fan Wang , Yi Yang

We present Mamoda2.5, a unified AR-Diffusion framework that seamlessly integrates multimodal understanding and generation within a single architecture. To efficiently enhance the model's generation capability, we equip the Diffusion…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Yangming Shi , Shixiang Zhu , Tao Shen , Zhimiao Yu , Dengsheng Chen , Taicai Chen , Yunfei Yang , Juan Zhou , Chen Cheng , Liang Ma , Xibin Wu , Benxuan Yan , Ge Li , Tuoyu Zhang , Dan Li , Chang Liu , Zhenbang Sun

Human-human motion generation is essential for understanding humans as social beings. Current methods fall into two main categories: single-person-based methods and separate modeling-based methods. To delve into this field, we abstract the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Yabiao Wang , Shuo Wang , Jiangning Zhang , Ke Fan , Jiafu Wu , Zhucun Xue , Yong Liu

Video-based world models hold significant potential for generating high-quality embodied manipulation data. However, current video generation methods struggle to achieve stable long-horizon generation: classical diffusion-based approaches…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Yu Shang , Lei Jin , Yiding Ma , Xin Zhang , Chen Gao , Wei Wu , Yong Li

In this work, we build a simple but strong baseline for sounding video generation. Given base diffusion models for audio and video, we integrate them with additional modules into a single model and train it to make the model jointly…

Machine Learning · Computer Science 2025-04-10 Masato Ishii , Akio Hayakawa , Takashi Shibuya , Yuki Mitsufuji

Generating realistic human motion is a central yet unsolved challenge in video generation. While reinforcement learning (RL)-based post-training has driven recent gains in general video quality, extending it to human motion remains…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Yidong Huang , Zun Wang , Han Lin , Dong-Ki Kim , Shayegan Omidshafiei , Jaehong Yoon , Jaemin Cho , Yue Zhang , Mohit Bansal

Recent advances in unified multimodal models indicate a clear trend towards comprehensive content generation. However, the auditory domain remains a significant challenge, with music and speech often developed in isolation, hindering…

Human Video Motion Transfer (HVMT) aims to, given an image of a source person, generate his/her video that imitates the motion of the driving person. Existing methods for HVMT mainly exploit Generative Adversarial Networks (GANs) to perform…

Computer Vision and Pattern Recognition · Computer Science 2022-09-02 Quanwei Yang , Xinchen Liu , Wu Liu , Hongtao Xie , Xiaoyan Gu , Lingyun Yu , Yongdong Zhang

Text-to-video diffusion models have enabled high-quality video synthesis, yet often fail to generate temporally coherent and physically plausible motion. A key reason is the models' insufficient understanding of complex motions that natural…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Aritra Bhowmik , Denis Korzhenkov , Cees G. M. Snoek , Amirhossein Habibian , Mohsen Ghafoorian

Most existing multimodality methods use separate backbones for autoregression-based discrete text generation and diffusion-based continuous visual generation, or the same backbone by discretizing the visual data to use autoregression for…

Computer Vision and Pattern Recognition · Computer Science 2024-09-25 Chuyang Zhao , Yuxing Song , Wenhao Wang , Haocheng Feng , Errui Ding , Yifan Sun , Xinyan Xiao , Jingdong Wang