English
Related papers

Related papers: MotionGPT: Finetuned LLMs Are General-Purpose Moti…

200 papers

Recent progress in large models has led to significant advances in unified multimodal generation and understanding. However, the development of models that unify motion-language generation and understanding remains largely underexplored.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Zekun Li , Sizhe An , Chengcheng Tang , Chuan Guo , Ivan Shugurov , Linguang Zhang , Amy Zhao , Srinath Sridhar , Lingling Tao , Abhay Mittal

Inspired by the recent success of LLMs, the field of human motion understanding has increasingly shifted toward developing large motion models. Despite some progress, current efforts remain far from achieving truly generalist models,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Ye Wang , Sipeng Zheng , Bin Cao , Qianshan Wei , Weishuai Zeng , Qin Jin , Zongqing Lu

We have recently seen tremendous progress in realistic text-to-motion generation. Yet, the existing methods often fail or produce implausible motions with unseen text inputs, which limits the applications. In this paper, we present OMG, a…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Han Liang , Jiacheng Bao , Ruichi Zhang , Sihan Ren , Yuecheng Xu , Sibei Yang , Xin Chen , Jingyi Yu , Lan Xu

Generating realistic human motions from textual descriptions has undergone significant advancements. However, existing methods often overlook specific body part movements and their timing. In this paper, we address this issue by enriching…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Bizhu Wu , Jinheng Xie , Meidan Ding , Zhe Kong , Jianfeng Ren , Ruibin Bai , Rong Qu , Linlin Shen

Humanoid robots, with their human-like embodiment, have the potential to integrate seamlessly into human environments. Critical to their coexistence and cooperation with humans is the ability to understand natural language communications…

Robotics · Computer Science 2024-10-17 Zhenyu Jiang , Yuqi Xie , Jinhan Li , Ye Yuan , Yifeng Zhu , Yuke Zhu

Programming a robotic is a complex task, as it demands the user to have a good command of specific programming languages and awareness of the robot's physical constraints. We propose a framework that simplifies robot deployment by allowing…

Human-Computer Interaction · Computer Science 2024-07-18 Cathy Mengying Fang , Krzysztof Zieliński , Pattie Maes , Joe Paradiso , Bruce Blumberg , Mikkel Baun Kjærgaard

Natural and lifelike locomotion remains a fundamental challenge for humanoid robots to interact with human society. However, previous methods either neglect motion naturalness or rely on unstable and ambiguous style rewards. In this paper,…

Robotics · Computer Science 2025-03-13 Haodong Zhang , Liang Zhang , Zhenghan Chen , Lu Chen , Yue Wang , Rong Xiong

We tackle the problem of generating long-term 3D human motion from multiple action labels. Two main previous approaches, such as action- and motion-conditioned methods, have limitations to solve this problem. The action-conditioned methods…

Computer Vision and Pattern Recognition · Computer Science 2023-02-20 Taeryung Lee , Gyeongsik Moon , Kyoung Mu Lee

While recently Multimodal Large Language Models (MM-LLMs) have made exciting strides, they mostly fall prey to the limitation of only input-side multimodal understanding, without the ability to produce content in multiple modalities. As we…

Artificial Intelligence · Computer Science 2024-06-26 Shengqiong Wu , Hao Fei , Leigang Qu , Wei Ji , Tat-Seng Chua

We present Lumina-mGPT, a family of multimodal autoregressive models capable of various vision and language tasks, particularly excelling in generating flexible photorealistic images from text descriptions. By initializing from multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-04-25 Dongyang Liu , Shitian Zhao , Le Zhuo , Weifeng Lin , Yi Xin , Xinyue Li , Qi Qin , Yu Qiao , Hongsheng Li , Peng Gao

We present PandaGPT, an approach to emPower large lANguage moDels with visual and Auditory instruction-following capabilities. Our pilot experiments show that PandaGPT can perform complex tasks such as detailed image description generation,…

Computation and Language · Computer Science 2023-05-29 Yixuan Su , Tian Lan , Huayang Li , Jialu Xu , Yan Wang , Deng Cai

In this paper, a deep learning-based model for 3D human motion generation from the text is proposed via gesture action classification and an autoregressive model. The model focuses on generating special gestures that express human thinking,…

Computer Vision and Pattern Recognition · Computer Science 2022-11-21 Gwantae Kim , Youngsuk Ryu , Junyeop Lee , David K. Han , Jeongmin Bae , Hanseok Ko

Text-driven human motion generation is an emerging task in animation and humanoid robot design. Existing algorithms directly generate the full sequence which is computationally expensive and prone to errors as it does not pay special…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Zichen Geng , Caren Han , Zeeshan Hayder , Jian Liu , Mubarak Shah , Ajmal Mian

Human motion generation is essential for fields such as animation, robotics, and virtual reality, requiring models that effectively capture motion dynamics from text descriptions. Existing approaches often rely on Contrastive Language-Image…

Computer Vision and Pattern Recognition · Computer Science 2025-05-19 Gabriel Maldonado , Armin Danesh Pazho , Ghazal Alinezhad Noghre , Vinit Katariya , Hamed Tabkhi

Conditional human motion generation is an important topic with many applications in virtual reality, gaming, and robotics. While prior works have focused on generating motion guided by text, music, or scenes, these typically result in…

Computer Vision and Pattern Recognition · Computer Science 2024-02-26 German Barquero , Sergio Escalera , Cristina Palmero

Learning behavior in legged robots presents a significant challenge due to its inherent instability and complex constraints. Recent research has proposed the use of a large language model (LLM) to generate reward functions in reinforcement…

Robotics · Computer Science 2025-07-01 Runhao Zeng , Dingjie Zhou , Qiwei Liang , Junlin Liu , Hui Li , Changxin Huang , Jianqiang Li , Xiping Hu , Fuchun Sun

In this paper, we tackle the problem of how to build and benchmark a large motion model (LMM). The ultimate goal of LMM is to serve as a foundation model for versatile motion-related tasks, e.g., human motion generation, with…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Liang Xu , Shaoyang Hua , Zili Lin , Yifan Liu , Feipeng Ma , Yichao Yan , Xin Jin , Xiaokang Yang , Wenjun Zeng

The advent of large language models, enabling flexibility through instruction-driven approaches, has revolutionized many traditional generative tasks, but large models for 3D data, particularly in comprehensively handling 3D shapes with…

Computer Vision and Pattern Recognition · Computer Science 2023-12-04 Fukun Yin , Xin Chen , Chi Zhang , Biao Jiang , Zibo Zhao , Jiayuan Fan , Gang Yu , Taihao Li , Tao Chen

Human motion generation is a significant pursuit in generative computer vision with widespread applications in film-making, video games, AR/VR, and human-robot interaction. Current methods mainly utilize either diffusion-based generative…

Computer Vision and Pattern Recognition · Computer Science 2025-02-03 Canxuan Gang

We introduce a novel method for real-time animation control and generation on rigged models using natural language input. First, we embed a large language model (LLM) in Unity to output structured texts that can be parsed into diverse and…

‹ Prev 1 4 5 6 7 8 10 Next ›