English
Related papers

Related papers: QuadFM: Foundational Text-Driven Quadruped Motion …

200 papers

This paper introduces MotionGlot, a model that can generate motion across multiple embodiments with different action dimensions, such as quadruped robots and human bodies. By leveraging the well-established training procedures commonly used…

Robotics · Computer Science 2025-05-02 Sudarshan Harithas , Srinath Sridhar

Quadruped animals are capable of exhibiting a diverse range of locomotion gaits. While progress has been made in demonstrating such gaits on robots, current methods rely on motion priors, dynamics models, or other forms of extensive manual…

Robotics · Computer Science 2024-02-23 David DeFazio , Yohei Hayamizu , Shiqi Zhang

Foundation Models (FM) have increasingly drawn the attention of researchers due to their scalability and generalization across diverse tasks. Inspired by the success of FMs and the principles that have driven advancements in Large Language…

Computer Vision and Pattern Recognition · Computer Science 2025-02-26 Mohammadreza Baharani , Ghazal Alinezhad Noghre , Armin Danesh Pazho , Gabriel Maldonado , Hamed Tabkhi

Generating realistic human motion from given action descriptions has experienced significant advancements because of the emerging requirement of digital humans. While recent works have achieved impressive results in generating motion…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Yaqi Zhang , Di Huang , Bin Liu , Shixiang Tang , Yan Lu , Lu Chen , Lei Bai , Qi Chu , Nenghai Yu , Wanli Ouyang

Our paper aims to generate diverse and realistic animal motion sequences from textual descriptions, without a large-scale animal text-motion dataset. While the task of text-driven human motion synthesis is already extensively studied and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Zhangsihao Yang , Mingyuan Zhou , Mengyi Shan , Bingbing Wen , Ziwei Xuan , Mitch Hill , Junjie Bai , Guo-Jun Qi , Yalin Wang

This paper introduces CognitiveDog, a pioneering development of quadruped robot with Large Multi-modal Model (LMM) that is capable of not only communicating with humans verbally but also physically interacting with the environment through…

We introduce Multimodal DuetDance (MDD), a diverse multimodal benchmark dataset designed for text-controlled and music-conditioned 3D duet dance motion generation. Our dataset comprises 620 minutes of high-quality motion capture data…

Graphics · Computer Science 2025-08-26 Prerit Gupta , Jason Alexander Fotso-Puepi , Zhengyuan Li , Jay Mehta , Aniket Bera

Text-to-motion (T2M) generation aims to create realistic human movements from text descriptions, with promising applications in animation and robotics. Despite recent progress, current T2M models perform poorly on unseen text descriptions…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Bin Cao , Sipeng Zheng , Hao Luo , Boyuan Li , Jing Liu , Zongqing Lu

Quadruped robots have become quite popular for their ability to adapt their locomotion to generic uneven terrains. For this reason, over time, several frameworks for quadrupedal locomotion have been proposed, but with little attention to…

Robotics · Computer Science 2025-04-30 Aristide Emanuele Casucci , Federico Nesti , Mauro Marinoni , Giorgio Buttazzo

Although existing text-to-motion (T2M) methods can produce realistic human motion from text description, it is still difficult to align the generated motion with the desired postures since using text alone is insufficient for precisely…

Computer Vision and Pattern Recognition · Computer Science 2025-04-03 Ling-An Zeng , Gaojie Wu , Ancong Wu , Jian-Fang Hu , Wei-Shi Zheng

Generating 3D human motion based on textual descriptions has been a research focus in recent years. It requires the generated motion to be diverse, natural, and conform to the textual description. Due to the complex spatio-temporal nature…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Chongyang Zhong , Lei Hu , Zihao Zhang , Shihong Xia

Text-to-motion (T2M) generation has broad applications in character animation, virtual avatars, and human-robot interaction. Existing methods typically generate pose trajectories or motion tokens directly from language, forcing a single…

Machine Learning · Computer Science 2026-05-29 Nikolay Shvetsov , Maksim Bobrin , Nazar Buzun , Dmitry V. Dylov

Human motion generation has advanced rapidly in recent years, yet the critical problem of creating spatially grounded, context-aware gestures has been largely overlooked. Existing models typically specialize either in descriptive motion…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Anna Deichler , Jim O'Regan , Teo Guichoux , David Johansson , Jonas Beskow

Text-to-motion generation has experienced remarkable progress in recent years. However, current approaches remain limited to synthesizing motion from short or general text prompts, primarily due to dataset constraints. This limitation…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Chuan Guo , Inwoo Hwang , Jian Wang , Bing Zhou

Quadruped robots face persistent challenges in achieving versatile locomotion due to limitations in reference motion data diversity. To address these challenges, we introduce an in-between motion generation based multi-style quadruped robot…

Robotics · Computer Science 2025-08-12 Yuanhao Chen , Liu Zhao , Ji Ma , Peng Lu

Large language models (LLMs) have demonstrated the potential to perform high-level planning. Yet, it remains a challenge for LLMs to comprehend low-level commands, such as joint angle targets or motor torques. This paper proposes an…

Robotics · Computer Science 2023-09-15 Yujin Tang , Wenhao Yu , Jie Tan , Heiga Zen , Aleksandra Faust , Tatsuya Harada

While quadruped robots usually have good stability and load capacity, bipedal robots offer a higher level of flexibility / adaptability to different tasks and environments. A multi-modal legged robot can take the best of both worlds. In…

Robotics · Computer Science 2022-02-25 Chen Yu , Andre Rosendo

Human communication is inherently multimodal, involving a combination of verbal and non-verbal cues such as speech, facial expressions, and body gestures. Modeling these behaviors is essential for understanding human interaction and for…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Changan Chen , Juze Zhang , Shrinidhi K. Lakshmikanth , Yusu Fang , Ruizhi Shao , Gordon Wetzstein , Li Fei-Fei , Ehsan Adeli

Mobility Foundation Models (MFMs) have advanced the modeling of human movement patterns, yet they face a ceiling due to limitations in data scale and semantic understanding. While Large Language Models (LLMs) offer powerful semantic…

Machine Learning · Computer Science 2026-04-14 Fanjin Meng , Yuan Yuan , Jingtao Ding , Jie Feng , Chonghua Han , Yong Li

Recent advances in large language models (LLMs) have enabled breakthroughs in many multimodal generation tasks, but a significant performance gap still exists in text-to-motion generation, where LLM-based methods lag far behind non-LLM…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Chuhao Jin , Haosen Li , Bingzi Zhang , Che Liu , Xiting Wang , Ruihua Song , Wenbing Huang , Ying Qin , Fuzheng Zhang , Di Zhang
‹ Prev 1 2 3 10 Next ›