English
Related papers

Related papers: UniAct: Unified Motion Generation and Action Strea…

200 papers

Recent advances in Multimodal Large Language Models (MLLMs) have driven rapid progress in Vision-Language-Action (VLA) models for robotic manipulation. Although effective in many scenarios, current approaches largely rely on explicit…

The evolution of autonomous systems in the context of human-robot interaction systems necessitates a synergy between the continuous perception of the environment and the potential actions to navigate or interact within it. We present…

Robotics · Computer Science 2025-09-18 Timothée Dhaussy , Bassam Jabaian , Fabrice Lefèvre

We present Human Motions with Objects (HUMOTO), a high-fidelity dataset of human-object interactions for motion generation, computer vision, and robotics applications. Featuring 735 sequences (7,875 seconds at 30 fps), HUMOTO captures…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Jiaxin Lu , Chun-Hao Paul Huang , Uttaran Bhattacharya , Qixing Huang , Yi Zhou

Enabling humanoid robots to follow free-form natural language commands is a critical step toward seamless human-robot interaction and general-purpose embodied AI. However, existing methods remain limited, often constrained to simple…

Robotics · Computer Science 2026-05-12 Zhirui Liu , Kaiyang Ji , Ke Yang , Yahao Fan , Jingyi Yu , Ye Shi , Jingya Wang

Humans are excellent at understanding language and vision to accomplish a wide range of tasks. In contrast, creating general instruction-following embodied agents remains a difficult challenge. Prior work that uses pure language-only models…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Hao Liu , Lisa Lee , Kimin Lee , Pieter Abbeel

Humanoid robots hold great potential for diverse interactions and daily service tasks within human-centered environments, necessitating controllers that seamlessly integrate precise locomotion with dexterous manipulation. However, most…

Robotics · Computer Science 2026-01-27 Xinru Cui , Linxi Feng , Yixuan Zhou , Haoqi Han , Zhe Liu , Hesheng Wang

From loco-motion to dextrous manipulation, humanoid robots have made remarkable strides in demonstrating complex full-body capabilities. However, the majority of current robot learning datasets and benchmarks mainly focus on stationary…

Large language models (LLMs) are, by design, inherently capable of multi-task learning: through a unified next-token prediction paradigm, they can naturally address a wide variety of downstream tasks. Prior work in the motion domain has…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Zeyu Ling , Bo Han , Shiyang Li , Jikang Cheng , Hongdeng Shen , Changqing Zou

The study of complex human interactions and group activities has become a focal point in human-centric computer vision. However, progress in related tasks is often hindered by the challenges of obtaining large-scale labeled datasets from…

Computer Vision and Pattern Recognition · Computer Science 2024-05-06 Che-Jui Chang , Danrui Li , Deep Patel , Parth Goel , Honglu Zhou , Seonghyeon Moon , Samuel S. Sohn , Sejong Yoon , Vladimir Pavlovic , Mubbasir Kapadia

End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generation models, limiting…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Gaojie Lin , Jianwen Jiang , Jiaqi Yang , Zerong Zheng , Chao Liang

In this paper, we consider unmanned aerial vehicles (UAVs) equipped with a visible light communication (VLC) access point and coordinated multipoint (CoMP) capability that allows users to connect to more than one UAV. UAVs can move in…

Signal Processing · Electrical Eng. & Systems 2021-12-06 Mohammad Reza Maleki , Mohammad Robat Mili , Mohammad Reza Javan , Nader Mokari , Eduard A. Jorswieck

Human motion driven control (HMDC) is an effective approach for generating natural and compelling robot motions while preserving high-level semantics. However, establishing the correspondence between humans and robots with different body…

Robotics · Computer Science 2023-10-02 Tianyu Li , Hyunyoung Jung , Matthew Gombolay , Yong Kwon Cho , Sehoon Ha

Whole-body humanoid teleoperation enables humans to remotely control humanoid robots, serving as both a real-time operational tool and a scalable engine for collecting demonstrations for autonomous learning. Despite recent advances,…

Robotics · Computer Science 2026-03-17 Yixuan Li , Le Ma , Yutang Lin , Yushi Du , Mengya Liu , Kaizhe Hu , Jieming Cui , Yixin Zhu , Wei Liang , Baoxiong Jia , Siyuan Huang

Unsupervised pre-training has shown great success in skeleton-based action understanding recently. Existing works typically train separate modality-specific models, then integrate the multi-modal information for action understanding by a…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Shengkai Sun , Daizong Liu , Jianfeng Dong , Xiaoye Qu , Junyu Gao , Xun Yang , Xun Wang , Meng Wang

Motion mimicking, i.e., encouraging the control policy to mimic human motion, facilitates the learning of complex tasks via reinforcement learning (RL) for humanoid robots. Although standard RL frameworks demonstrate impressive locomotion…

Robotics · Computer Science 2026-03-10 Ludwig Chee-Ying Tay , I-Chia Chang , Yan Gu

Understanding action correspondence between humans and robots is essential for evaluating alignment in decision-making, particularly in human-robot collaboration and imitation learning within unstructured environments. We propose a…

Robotics · Computer Science 2025-04-17 Azizul Zahid , Jie Fan , Farong Wang , Ashton Dy , Sai Swaminathan , Fei Liu

A person's movement or relative positioning can be effectively captured by different types of sensors and corresponding sensor output can be utilized in various manipulative techniques for the classification of different human activities.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Utsab Saha , Sawradip Saha , Tahmid Kabir , Shaikh Anowarul Fattah , Mohammad Saquib

Advancements in Multimodal Large Language Models (MLLMs) have improved human motion understanding. However, these models remain constrained by their "instruct-only" nature, lacking interactivity and adaptability for diverse analytical…

Artificial Intelligence · Computer Science 2025-02-28 Lei Li , Sen Jia , Jianhao Wang , Zhaochong An , Jiaang Li , Jenq-Neng Hwang , Serge Belongie

Continuum robots, which often rely on interdisciplinary and multimedia collaborations, have been increasingly recognized for their potential to revolutionize the field of human-computer interaction (HCI) in varied applications due to their…

Robotics · Computer Science 2024-10-08 Po-Yu Hsieh , June-Hao Hou

In recent years, reinforcement learning and imitation learning have shown great potential for controlling humanoid robots' motion. However, these methods typically create simulation environments and rewards for specific tasks, resulting in…

Robotics · Computer Science 2024-08-01 Jingkai Sun , Qiang Zhang , Yiqun Duan , Xiaoyang Jiang , Chong Cheng , Renjing Xu