English
Related papers

Related papers: MotionScript: Natural Language Descriptions for Ex…

200 papers

We introduce Unimotion, the first unified multi-task human motion model capable of both flexible motion control and frame-level motion understanding. While existing works control avatar motion with global text conditioning, or with…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Chuqiao Li , Julian Chibane , Yannan He , Naama Pearl , Andreas Geiger , Gerard Pons-moll

Text-driven human motion synthesis has showcased its potential for revolutionizing motion design in the movie and game industry. Existing methods often rely on 3D motion capture data, which requires special setups, resulting in high costs…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Ruoxi Guo , Huaijin Pi , Zehong Shen , Qing Shuai , Zechen Hu , Zhumei Wang , Yajiao Dong , Ruizhen Hu , Taku Komura , Sida Peng , Xiaowei Zhou

This paper introduces a novel dataset construction pipeline that samples pairs of frames from videos and uses multimodal large language models (MLLMs) to generate editing instructions for training instruction-based image manipulation…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Mingdeng Cao , Xuaner Zhang , Yinqiang Zheng , Zhihao Xia

Text-based 3D human motion editing is a critical yet challenging task in computer vision and graphics. While training-free approaches have been explored, the recent release of the MotionFix dataset, which includes source-text-motion…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Zhengyuan Li , Kai Cheng , Anindita Ghosh , Uttaran Bhattacharya , Liangyan Gui , Aniket Bera

This report reviews recent advancements in human motion prediction, reconstruction, and generation. Human motion prediction focuses on forecasting future poses and movements from historical data, addressing challenges like nonlinear…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Canxuan Gang , Yiran Wang

We introduce Programmatic Motion Concepts, a hierarchical motion representation for human actions that captures both low-level motion and high-level description as motion concepts. This representation enables human motion description,…

Computer Vision and Pattern Recognition · Computer Science 2022-06-28 Sumith Kulal , Jiayuan Mao , Alex Aiken , Jiajun Wu

Recent advances in world models have demonstrated strong capabilities in simulating physical reality, making them an increasingly important foundation for embodied intelligence. For UAV agents in particular, accurate prediction of complex…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Zile Guo , Zhan Chen , Enze Zhu , Kan Wei , Yongkang Zou , Xiaoxuan Liu , Lei Wang

Human-centred systems require an understanding of human actions in the physical world. Temporally extended sequences of actions are intentional and structured, yet existing methods for recognising what actions are performed often do not…

Artificial Intelligence · Computer Science 2026-04-21 Rimvydas Rubavicius , Manisha Dubey , N. Siddharth , Subramanian Ramamoorthy

We present a system that allows users to visualize complex human motion via 3D motion sculptures---a representation that conveys the 3D structure swept by a human body as it moves through space. Given an input video, our system computes the…

Human-Computer Interaction · Computer Science 2019-01-03 Xiuming Zhang , Tali Dekel , Tianfan Xue , Andrew Owens , Qiurui He , Jiajun Wu , Stefanie Mueller , William T. Freeman

In this paper, a deep learning-based model for 3D human motion generation from the text is proposed via gesture action classification and an autoregressive model. The model focuses on generating special gestures that express human thinking,…

Computer Vision and Pattern Recognition · Computer Science 2022-11-21 Gwantae Kim , Youngsuk Ryu , Junyeop Lee , David K. Han , Jeongmin Bae , Hanseok Ko

If robots are to work effectively alongside people, they must be able to interpret natural language references to objects in their 3D environment. Understanding 3D referring expressions is challenging -- it requires the ability to both…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Jiading Fang , Xiangshan Tan , Shengjie Lin , Igor Vasiljevic , Vitor Guizilini , Hongyuan Mei , Rares Ambrus , Gregory Shakhnarovich , Matthew R Walter

This paper introduces MotionGlot, a model that can generate motion across multiple embodiments with different action dimensions, such as quadruped robots and human bodies. By leveraging the well-established training procedures commonly used…

Robotics · Computer Science 2025-05-02 Sudarshan Harithas , Srinath Sridhar

Text-to-motion models that generate sequences of human poses from textual descriptions are garnering significant attention. However, due to data scarcity, the range of motions these models can produce is still limited. For instance, current…

Machine Learning · Computer Science 2024-10-14 Clayton Leite , Yu Xiao

Motions of virtual characters in movies or video games are typically generated by recording actors using motion capturing methods. Animations generated this way often need postprocessing, such as improving the periodicity of cyclic…

Computer Vision and Pattern Recognition · Computer Science 2015-02-27 Martin Bauer , Markus Eslitzbichler , Markus Grasmair

We introduce the Cross Human Motion Diffusion Model (CrossDiff), a novel approach for generating high-quality human motion based on textual descriptions. Our method integrates 3D and 2D information using a shared transformer network within…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Zeping Ren , Shaoli Huang , Xiu Li

Text-driven human motion generation is a multimodal task that synthesizes human motion sequences conditioned on natural language. It requires the model to satisfy textual descriptions under varying conditional inputs, while generating…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Xingyu Chen

Generating 3D human motions from textual descriptions is an important research problem with broad applications in video games, virtual reality, and augmented reality. Recent methods align the textual description with human motion at the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-25 Bowen Dang , Lin Wu , Xiaohang Yang , Zheng Yuan , Zhixiang Chen

Humanoid robots, with their human-like embodiment, have the potential to integrate seamlessly into human environments. Critical to their coexistence and cooperation with humans is the ability to understand natural language communications…

Robotics · Computer Science 2024-10-17 Zhenyu Jiang , Yuqi Xie , Jinhan Li , Ye Yuan , Yifeng Zhu , Yuke Zhu

3D Human Motion Indexing and Retrieval is an interesting problem due to the rise of several data-driven applications aimed at analyzing and/or re-utilizing 3D human skeletal data, such as data-driven animation, analysis of sports…

Computer Vision and Pattern Recognition · Computer Science 2019-12-11 Neeraj Battan , Abbhinav Venkat , Avinash Sharma

Despite advancements in Multimodal Large Language Models (MLLMs), their proficiency in fine-grained video motion understanding remains critically limited. They often lack inter-frame differencing and tend to average or ignore subtle visual…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Yipeng Du , Tiehan Fan , Kepan Nan , Rui Xie , Penghao Zhou , Xiang Li , Jian Yang , Zhenheng Yang , Ying Tai