English
Related papers

Related papers: UniAct: Unified Motion Generation and Action Strea…

200 papers

Training on diverse, internet-scale data is a key factor in the success of recent large foundation models. Yet, using the same recipe for building embodied agents has faced noticeable difficulties. Despite the availability of many…

A major challenge in humanoid robotics is designing a unified interface for commanding diverse whole-body behaviors, from precise footstep sequences to partial-body mimicry and joystick teleoperation. We introduce the Masked Humanoid…

Robotics · Computer Science 2026-04-23 Pranay Dugar , Aayam Shrestha , Fangzhou Yu , Bart van Marum , Alan Fern

Loco-Manipulation for humanoid robots aims to enable robots to integrate mobility with upper-body tracking capabilities. Most existing approaches adopt hierarchical architectures that decompose control into isolated upper-body…

Robotics · Computer Science 2026-03-03 Wandong Sun , Luying Feng , Baoshi Cao , Yang Liu , Yaochu Jin , Zongwu Xie

This paper introduces OmniMotion-X, a versatile multimodal framework for whole-body human motion generation, leveraging an autoregressive diffusion transformer in a unified sequence-to-sequence manner. OmniMotion-X efficiently supports…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Guowei Xu , Yuxuan Bian , Ailing Zeng , Mingyi Shi , Shaoli Huang , Wen Li , Lixin Duan , Qiang Xu

We present Action Agent, a two-stage framework that unifies agentic navigation video generation with flow-constrained diffusion control for multi-embodiment robot navigation. In Stage I, a large language model (LLM) acts as an orchestration…

Robotics · Computer Science 2026-05-05 Jeffrin Sam , Nguyen Khang , Yara Mahmoud , Miguel Altamirano Cabrera , Dzmitry Tsetserukou

Achieving expressive and generalizable whole-body motion control is essential for deploying humanoid robots in real-world environments. In this work, we propose UniTracker, a three-stage training framework that enables robust and scalable…

Real-time whole-body teleoperation is a critical method for humanoid robots to perform complex tasks in unstructured environments. However, developing a unified controller that robustly supports diverse human motions remains a significant…

Robotics · Computer Science 2026-05-14 Jie Li , Bing Tang , Feng Wu

Interactive humanoid video generation aims to synthesize lifelike visual agents that can engage with humans through continuous and responsive video. Despite recent advances in video synthesis, existing methods often grapple with the…

Enabling humanoid robots to perform agile and adaptive interactive tasks has long been a core challenge in robotics. Current approaches are bottlenecked by either the scarcity of realistic interaction data or the need for meticulous,…

Robotics · Computer Science 2026-02-03 Yinhuai Wang , Qihan Zhao , Yuen Fui Lau , Runyi Yu , Hok Wai Tsui , Qifeng Chen , Jingbo Wang , Jiangmiao Pang , Ping Tan

Equipping humanoid robots with versatile interaction skills typically requires either extensive policy training or explicit human-to-robot motion retargeting. However, learning-based policies face prohibitive data collection costs.…

Scaling humanoid foundation models is bottlenecked by the scarcity of robotic data. While massive egocentric human data offers a scalable alternative, bridging the cross-embodiment chasm remains a fundamental challenge due to kinematic…

Robotics · Computer Science 2026-04-22 Boyu Chen , Yi Chen , Lu Qiu , Jerry Bai , Yuying Ge , Yixiao Ge

Simulated humanoids are an appealing research domain due to their physical capabilities. Nonetheless, they are also challenging to control, as a policy must drive an unstable, discontinuous, and high-dimensional physical system. One widely…

We present a scalable framework for cross-embodiment humanoid robot control by learning a shared latent representation that unifies motion across humans and diverse humanoid platforms, including single-arm, dual-arm, and legged humanoid…

Robotics · Computer Science 2026-01-23 Yashuai Yan , Dongheui Lee

We tackle the problem of generating long-term 3D human motion from multiple action labels. Two main previous approaches, such as action- and motion-conditioned methods, have limitations to solve this problem. The action-conditioned methods…

Computer Vision and Pattern Recognition · Computer Science 2023-02-20 Taeryung Lee , Gyeongsik Moon , Kyoung Mu Lee

Large language models (LLMs) have unified diverse linguistic tasks within a single framework, yet such unification remains unexplored in human motion generation. Existing methods are confined to isolated tasks, limiting flexibility for…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Wendong Bu , Kaihang Pan , Yuze Lin , Jiacheng Li , Kai Shen , Wenqiao Zhang , Juncheng Li , Jun Xiao , Siliang Tang

This work focuses on generating realistic, physically-based human behaviors from multi-modal inputs, which may only partially specify the desired motion. For example, the input may come from a VR controller providing arm motion and body…

Robotics · Computer Science 2025-02-11 Aayam Shrestha , Pan Liu , German Ros , Kai Yuan , Alan Fern

Learning a general humanoid whole-body controller is challenging because practical reference motions can exhibit noise and inconsistencies after being transferred to the robot domain, and local defects may be amplified by closed-loop…

Robotics · Computer Science 2026-02-02 Yubiao Ma , Han Yu , Jiayin Xie , Changtai Lv , Qiang Luo , Chi Zhang , Yunpeng Yin , Boyang Xing , Xuemei Ren , Dongdong Zheng

A generalist robotic policy needs both semantic understanding for task planning and the ability to interact with the environment through predictive capabilities. To tackle this, we present MM-ACT, a unified Vision-Language-Action (VLA)…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Haotian Liang , Xinyi Chen , Bin Wang , Mingkang Chen , Yitian Liu , Yuhao Zhang , Zanxin Chen , Tianshuo Yang , Yilun Chen , Jiangmiao Pang , Dong Liu , Xiaokang Yang , Yao Mu , Wenqi Shao , Ping Luo

Whole-body humanoid motion represents a fundamental challenge in robotics, requiring balance, coordination, and adaptability to enable human-like behaviors. However, existing methods typically require multiple training samples per motion,…

Enabling robust whole-body humanoid-object interaction (HOI) remains challenging due to motion data scarcity and the contact-rich nature. We present HDMI (HumanoiD iMitation for Interaction), a simple and general framework that learns…

Robotics · Computer Science 2025-09-30 Haoyang Weng , Yitang Li , Nikhil Sobanbabu , Zihan Wang , Zhengyi Luo , Tairan He , Deva Ramanan , Guanya Shi
‹ Prev 1 2 3 10 Next ›