English
Related papers

Related papers: TEMOS: Generating diverse human motions from textu…

200 papers

Recent advances in text-driven human motion generation enable models to synthesize realistic motion sequences from natural language descriptions. However, most existing approaches assume identity-neutral motion and generate movements using…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Wenqi Jia , Zekun Li , Abhay Mittal , Chengcheng Tang , Chuan Guo , Lezi Wang , James Matthew Rehg , Lingling Tao , Size An

Recently, text-to-motion models have opened new possibilities for creating realistic human motion with greater efficiency and flexibility. However, aligning motion generation with event-level textual descriptions presents unique challenges…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Haonan Han , Xiangzuo Wu , Huan Liao , Zunnan Xu , Zhongyuan Hu , Ronghui Li , Yachao Zhang , Xiu Li

Success in generative modeling across language, image, and video demonstrates that large, well-curated datasets are the key driver for building capable models. 3D Human motion, however, has lagged behind, constrained by an unsatisfying…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Jiahao Zhang , Joseph Liu , Young-Yoon Lee , Seonghyeon Moon , Victor Zordan , Guy Tevet , Karen Liu , Stephen Gould , Oren Jacob , Haomiao Jiang , Mubbasir Kapadia , Yizhak Ben-Shabat

Human-motion generation is a long-standing challenging task due to the requirement of accurately modeling complex and diverse dynamic patterns. Most existing methods adopt sequence models such as RNN to directly model transitions in the…

Computer Vision and Pattern Recognition · Computer Science 2019-12-24 Zhenyi Wang , Ping Yu , Yang Zhao , Ruiyi Zhang , Yufan Zhou , Junsong Yuan , Changyou Chen

We present a novel alignment-before-generation approach to tackle the challenging task of generating general 3D shapes based on 2D images or texts. Directly learning a conditional generative model from images or texts to 3D shapes is prone…

Computer Vision and Pattern Recognition · Computer Science 2023-07-04 Zibo Zhao , Wen Liu , Xin Chen , Xianfang Zeng , Rui Wang , Pei Cheng , Bin Fu , Tao Chen , Gang Yu , Shenghua Gao

Generating human motion from text has been dominated by denoising motion models either through diffusion or generative masking process. However, these models face great limitations in usability by requiring prior knowledge of the motion…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Ekkasit Pinyoanuntapong , Muhammad Usama Saleem , Pu Wang , Minwoo Lee , Srijan Das , Chen Chen

Generating human videos with realistic and controllable motions is a challenging task. While existing methods can generate visually compelling videos, they lack separate control over four key video elements: foreground subject, background…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Jingyun Liang , Jingkai Zhou , Shikai Li , Chenjie Cao , Lei Sun , Yichen Qian , Weihua Chen , Fan Wang

Learning 3D human motion from 2D inputs is a fundamental task in the realms of computer vision and computer graphics. Many previous methods grapple with this inherently ambiguous task by introducing motion priors into the learning process.…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Shuaiying Hou , Hongyu Tao , Junheng Fang , Changqing Zou , Hujun Bao , Weiwei Xu

Current techniques face difficulties in generating motions from intricate semantic descriptions, primarily due to insufficient semantic annotations in datasets and weak contextual understanding. To address these issues, we present…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Xin He , Shaoli Huang , Xiaohang Zhan , Chao Weng , Ying Shan

The automatic generation of controllable co-speech gestures has recently gained growing attention. While existing systems typically achieve gesture control through predefined categorical labels or implicit pseudo-labels derived from motion…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Bohong Chen , Yumeng Li , Youyi Zheng , Yao-Xiang Ding , Kun Zhou

Text-to-motion generation requires not only grounding local actions in language but also seamlessly blending these individual actions to synthesize diverse and realistic global motions. However, existing motion generation methods primarily…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Peng Jin , Hao Li , Zesen Cheng , Kehan Li , Runyi Yu , Chang Liu , Xiangyang Ji , Li Yuan , Jie Chen

Synthesizing text-driven 3D human motion within realistic scenes requires learning both semantic intent ("walk to the couch") and physical feasibility (e.g., avoiding collisions). Current methods use generative frameworks that…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Anindita Ghosh , Vladislav Golyanik , Taku Komura , Philipp Slusallek , Christian Theobalt , Rishabh Dabral

Modeling human-human interactions from text remains challenging because it requires not only realistic individual dynamics but also precise, text-consistent spatiotemporal coupling between agents. Currently, progress is hindered by 1)…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Qingxuan Wu , Zhiyang Dou , Chuan Guo , Yiming Huang , Qiao Feng , Bing Zhou , Jian Wang , Lingjie Liu

Though the advancement of pre-trained large language models unfolds, the exploration of building a unified model for language and other multi-modal data, such as motion, remains challenging and untouched so far. Fortunately, human motion…

Computer Vision and Pattern Recognition · Computer Science 2023-07-21 Biao Jiang , Xin Chen , Wen Liu , Jingyi Yu , Gang Yu , Tao Chen

Human video generation is a dynamic and rapidly evolving task that aims to synthesize 2D human body video sequences with generative models given control conditions such as text, audio, and pose. With the potential for wide-ranging…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Wentao Lei , Jinting Wang , Fengji Ma , Guanjie Huang , Li Liu

While recent text-to-video models excel at generating diverse scenes, they struggle with precise motion control, particularly for complex, multi-subject motions. Although methods for single-motion customization have been developed to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Youcan Xu , Zhen Wang , Jiaxin Shi , Kexin Li , Feifei Shao , Jun Xiao , Yi Yang , Jun Yu , Long Chen

Automatic gesture synthesis from speech is a topic that has attracted researchers for applications in remote communication, video games and Metaverse. Learning the mapping between speech and 3D full-body gestures is difficult due to the…

Computer Vision and Pattern Recognition · Computer Science 2023-10-12 Kunkun Pang , Dafei Qin , Yingruo Fan , Julian Habekost , Takaaki Shiratori , Junichi Yamagishi , Taku Komura

Human motion modeling is important for many modern graphics applications, which typically require professional skills. In order to remove the skill barriers for laymen, recent motion generation methods can directly generate human motions…

Computer Vision and Pattern Recognition · Computer Science 2022-09-01 Mingyuan Zhang , Zhongang Cai , Liang Pan , Fangzhou Hong , Xinying Guo , Lei Yang , Ziwei Liu

Recent advances in interactive technologies have highlighted the prominence of audio signals for semantic encoding. This paper explores a new task, where audio signals are used as conditioning inputs to generate motions that align with the…

Sound · Computer Science 2025-05-30 Zi-An Wang , Shihao Zou , Shiyao Yu , Mingyuan Zhang , Chao Dong

Embodied human communication encompasses both verbal (speech) and non-verbal information (e.g., gesture and head movements). Recent advances in machine learning have substantially improved the technologies for generating synthetic versions…

Machine Learning · Computer Science 2021-01-15 Simon Alexanderson , Éva Székely , Gustav Eje Henter , Taras Kucherenko , Jonas Beskow