English
Related papers

Related papers: HumanTOMATO: Text-aligned Whole-body Motion Genera…

200 papers

Recent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where text describes the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Bizhu Wu , Jinheng Xie , Keming Shen , Zhe Kong , Jianfeng Ren , Ruibin Bai , Rong Qu , Linlin Shen

Generating videos of complex human motions such as flips, cartwheels, and martial arts remains challenging for current video diffusion models. Text-only conditioning is temporally ambiguous for fine-grained motion control, while explicit…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Ashkan Taghipour , Morteza Ghahremani , Zinuo Li , Hamid Laga , Farid Boussaid , Mohammed Bennamoun

Text-to-motion generation is driven by learning motion representations for semantic alignment with language. Existing methods rely on either continuous or discrete motion representations. However, continuous representations entangle…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Dawei Guan , Di Yang , Chengjie Jin , Jiangtao Wang

We introduce an approach for augmenting text-to-video generation models with customized motions, extending their capabilities beyond the motions depicted in the original training data. By leveraging a few video samples demonstrating…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Joanna Materzynska , Josef Sivic , Eli Shechtman , Antonio Torralba , Richard Zhang , Bryan Russell

Language-guided human motion synthesis has been a challenging task due to the inherent complexity and diversity of human behaviors. Previous methods face limitations in generalization to novel actions, often resulting in unrealistic or…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Yuanhao Zhai , Mingzhen Huang , Tianyu Luan , Lu Dong , Ifeoma Nwogu , Siwei Lyu , David Doermann , Junsong Yuan

Whole-body multi-modal human motion generation poses two primary challenges: creating an effective motion generation mechanism and integrating various modalities, such as text, speech, and music, into a cohesive framework. Unlike previous…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Zhe Li , Weihao Yuan , Weichao Shen , Siyu Zhu , Zilong Dong , Chang Xu

Existing humanoid control systems often rely on teleoperation or modular generation pipelines that separate language understanding from physical execution. However, the former is entirely human-driven, and the latter lacks tight alignment…

Robotics · Computer Science 2025-11-25 Yuxuan Wang , Haobin Jiang , Shiqing Yao , Ziluo Ding , Zongqing Lu

Recent advances in humanoid whole-body motion tracking have enabled the execution of diverse and highly coordinated motions on real hardware. However, existing controllers are commonly driven either by predefined motion trajectories, which…

Robotics · Computer Science 2026-02-10 Weiji Xie , Jiakun Zheng , Jinrui Han , Jiyuan Shi , Weinan Zhang , Chenjia Bai , Xuelong Li

Training language-conditioned whole-body controllers for humanoid robots demands large-scale motion-language datasets. Existing approaches based on motion capture are costly and limited in diversity, while text-to-motion generative models…

Robotics · Computer Science 2026-04-20 Jianuo Cao , Yuxin Chen , Masayoshi Tomizuka

In filmmaking, directors typically allow actors to perform freely based on the script before providing specific guidance on how to present key actions. AI-generated content faces similar requirements, where users not only need automatic…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Zheng Qin , Ruobing Zheng , Yabing Wang , Tianqi Li , Zixin Zhu , Sanping Zhou , Ming Yang , Le Wang

Recent advancements in diffusion models have significantly improved the realism and generalizability of character-driven animation, enabling the synthesis of high-quality motion from just a single RGB image and a set of driving poses.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Alireza Javanmardi , Pragati Jaiswal , Tewodros Amberbir Habtegebrial , Christen Millerdurai , Shaoxiang Wang , Alain Pagani , Didier Stricker

In this paper, we introduce RoleMotion, a large-scale human motion dataset that encompasses a wealth of role-playing and functional motion data tailored to fit various specific scenes. Existing text datasets are mainly constructed…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Junran Peng , Yiheng Huang , Silei Shen , Zeji Wei , Jingwei Yang , Baojie Wang , Yonghao He , Chuanchen Luo , Man Zhang , Xucheng Yin , Wei Sui

While previous approaches to 3D human motion generation have achieved notable success, they often rely on extensive training and are limited to specific tasks. To address these challenges, we introduce Motion-Agent, an efficient…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Qi Wu , Yubo Zhao , Yifan Wang , Xinhang Liu , Yu-Wing Tai , Chi-Keung Tang

Body and face motion play an integral role in communication. They convey crucial information on the participants. Advances in generative modeling and multi-modal learning have enabled motion generation from signals such as speech,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Lownish Rai Sookha , Nikhil Pakhale , Mudasir Ganaie , Abhinav Dhall

Text-conditioned motion synthesis has made remarkable progress with the emergence of diffusion models. However, the majority of these motion diffusion models are primarily designed for a single character and overlook multi-human…

Computer Vision and Pattern Recognition · Computer Science 2024-11-22 Zhenzhi Wang , Jingbo Wang , Yixuan Li , Dahua Lin , Bo Dai

3D Human motion generation is pivotal across film, animation, gaming, and embodied intelligence. Traditional 3D motion synthesis relies on costly motion capture, while recent work shows that 2D videos provide rich, temporally coherent…

Graphics · Computer Science 2026-05-20 Yi-Yang Zhang , Tengjiao Sun , Pengcheng Fang , Deng-Bao Wang , Xiaohao Cai , Min-Ling Zhang , Hansung Kim

We study a challenging task: text-to-motion synthesis, aiming to generate motions that align with textual descriptions and exhibit coordinated movements. Currently, the part-based methods introduce part partition into the motion synthesis…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Qiran Zou , Shangyuan Yuan , Shian Du , Yu Wang , Chang Liu , Yi Xu , Jie Chen , Xiangyang Ji

With the rapid advancement of intelligent transportation systems, text-driven image generation and editing techniques have demonstrated significant potential in providing rich, controllable visual scene data for applications such as traffic…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Feng Lv , Haoxuan Feng , Zilu Zhang , Chunlong Xia , Yanfeng Li

The synthesis of human grasping has numerous applications including AR/VR, video games and robotics. While methods have been proposed to generate realistic hand-object interaction for object grasping and manipulation, these typically only…

Computer Vision and Pattern Recognition · Computer Science 2022-08-18 Yan Wu , Jiahao Wang , Yan Zhang , Siwei Zhang , Otmar Hilliges , Fisher Yu , Siyu Tang

Recent advances in diffusion-based video generation have enabled photo-realistic short clips, but current methods still struggle to achieve multi-modal consistency when jointly generating whole-body motion and natural speech. Current…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Xinhan Di , Kristin Qi , Pengqian Yu
‹ Prev 1 4 5 6 7 8 10 Next ›