English
Related papers

Related papers: QuadFM: Foundational Text-Driven Quadruped Motion …

200 papers

Our research presents a novel motion generation framework designed to produce whole-body motion sequences conditioned on multiple modalities simultaneously, specifically text and audio inputs. Leveraging Vector Quantized Variational…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Sohan Anisetty , James Hays

In this work, we propose TextIM, a novel framework for synthesizing TEXT-driven human Interactive Motions, with a focus on the precise alignment of part-level semantics. Existing methods often overlook the critical roles of interactive body…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Siyuan Fan , Bo Du , Xiantao Cai , Bo Peng , Longling Sun

We introduce KumoRFM-2, the next iteration of a pre-trained foundation model for relational data. KumoRFM-2 supports in-context learning as well as fine-tuning and is applicable to a wide range of predictive tasks. In contrast to tabular…

Machine Learning · Computer Science 2026-04-15 Valter Hudovernik , Federico López , Vid Kocijan , Akihiro Nitta , Jan Eric Lenssen , Jure Leskovec , Matthias Fey

Wearable accelerometers are used for a wide range of applications, such as gesture recognition, gait analysis, and sports monitoring. Yet most existing foundation models focus primarily on classifying common daily activities such as…

Machine Learning · Computer Science 2025-09-29 Junyong Park , Oron Levy , Rebecca Adaimi , Asaf Liberman , Gierad Laput , Abdelkareem Bedri

Generating high-quality cartoon animations multimodal control is challenging due to the complexity of non-human characters, stylistically diverse motions and fine-grained emotions. There is a huge domain gap between real-world videos and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Shuolin Xu , Bingyuan Wang , Zeyu Cai , Fangteng Fu , Yue Ma , Tongyi Lee , Hongchuan Yu , Zeyu Wang

Human motion modeling is important for many modern graphics applications, which typically require professional skills. In order to remove the skill barriers for laymen, recent motion generation methods can directly generate human motions…

Computer Vision and Pattern Recognition · Computer Science 2022-09-01 Mingyuan Zhang , Zhongang Cai , Liang Pan , Fangzhou Hong , Xinying Guo , Lei Yang , Ziwei Liu

Human video generation is becoming an increasingly important task with broad applications in graphics, entertainment, and embodied AI. Despite the rapid progress of video diffusion models (VDMs), their use for general-purpose human video…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Hyelin Nam , Hyojun Go , Byeongjun Park , Byung-Hoon Kim , Hyungjin Chung

Foundational models (FMs), pretrained on extensive datasets using self-supervised techniques, are capable of learning generalized patterns from large amounts of data. This reduces the need for extensive labeled datasets for each new task,…

Machine Learning · Computer Science 2024-06-19 Quan M. Tran , Suong N. Hoang , Lam M. Nguyen , Dzung Phan , Hoang Thanh Lam

In this study, we aim to initiate the development of Radiology Foundation Model, termed as RadFM. We consider the construction of foundational models from three perspectives, namely, dataset construction, model design, and thorough…

Computer Vision and Pattern Recognition · Computer Science 2023-11-17 Chaoyi Wu , Xiaoman Zhang , Ya Zhang , Yanfeng Wang , Weidi Xie

Diffusion models are increasingly used for robot learning, but current designs face a clear trade-off. Action-chunking diffusion policies like ManiCM are fast to run, yet they only predict short segments of motion. This makes them reactive,…

Robotics · Computer Science 2026-03-27 Xirui Shi , Arya Ebrahimi , Yi Hu , Jun Jin

Achieving realistic, vivid, and human-like synthesized conversational gestures conditioned on multi-modal data is still an unsolved problem due to the lack of available datasets, models and standard evaluation metrics. To address this, we…

Computer Vision and Pattern Recognition · Computer Science 2022-09-21 Haiyang Liu , Zihao Zhu , Naoya Iwamoto , Yichen Peng , Zhengqing Li , You Zhou , Elif Bozkurt , Bo Zheng

Recently reinforcement learning (RL) has emerged as a promising approach for quadrupedal locomotion, which can save the manual effort in conventional approaches such as designing skill-specific controllers. However, due to the complex…

Robotics · Computer Science 2021-09-17 Haojie Shi , Bo Zhou , Hongsheng Zeng , Fan Wang , Yueqiang Dong , Jiangyong Li , Kang Wang , Hao Tian , Max Q. -H. Meng

Whole-body multimodal motion generation, controlled by text, speech, or music, has numerous applications including video generation and character animation. However, employing a unified model to achieve various generation tasks with…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Yuxuan Bian , Ailing Zeng , Xuan Ju , Xian Liu , Zhaoyang Zhang , Wei Liu , Qiang Xu

We address the challenging problem of fine-grained text-driven human motion generation. Existing works generate imprecise motions that fail to accurately capture relationships specified in text due to: (1) lack of effective text parsing for…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Yin Wang , Mu Li , Jiapeng Liu , Zhiying Leng , Frederick W. B. Li , Ziyao Zhang , Xiaohui Liang

This work targets a novel text-driven whole-body motion generation task, which takes a given textual description as input and aims at generating high-quality, diverse, and coherent facial expressions, hand gestures, and body motions…

Computer Vision and Pattern Recognition · Computer Science 2023-10-20 Shunlin Lu , Ling-Hao Chen , Ailing Zeng , Jing Lin , Ruimao Zhang , Lei Zhang , Heung-Yeung Shum

Designing agile locomotion for quadruped robots often requires extensive expertise and tedious manual tuning. In this paper, we present a system to automate this process by leveraging deep reinforcement learning techniques. Our system can…

Unified physics-based humanoid controllers are pivotal for robotics and character animation, yet models that excel on gentle, everyday motions still stumble on explosive actions, hampering real-world deployment. We bridge this gap with FARM…

Robotics · Computer Science 2025-08-28 Tan Jing , Shiting Chen , Yangfan Li , Weisheng Xu , Renjing Xu

Humans and animals are believed to use a very minimal set of trajectories to perform a wide variety of tasks including walking. Our main objective in this paper is two fold 1) Obtain an effective tool to realize these basic motion patterns…

Generative masked transformers have demonstrated remarkable success across various content generation tasks, primarily due to their ability to effectively model large-scale dataset distributions with high consistency. However, in the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Yilin Wang , Chuan Guo , Yuxuan Mu , Muhammad Gohar Javed , Xinxin Zuo , Juwei Lu , Hai Jiang , Li Cheng

This work presents a motion retargeting approach for legged robots, aimed at transferring the dynamic and agile movements to robots from source motions. In particular, we guide the imitation learning procedures by transferring motions from…

Robotics · Computer Science 2025-07-25 Taerim Yoon , Dongho Kang , Seungmin Kim , Jin Cheng , Minsung Ahn , Stelian Coros , Sungjoon Choi