English
Related papers

Related papers: Learning Human-Intention Priors from Large-Scale H…

200 papers

Humans inherently possess generalizable visual representations that empower them to efficiently explore and interact with the environments in manipulation tasks. We advocate that such a representation automatically arises from…

Despite significant progress in robotic systems for operation within human-centric environments, existing models still heavily rely on explicit human commands to identify and manipulate specific objects. This limits their effectiveness in…

Robotics · Computer Science 2024-10-16 Shiyu Jin , Jinxuan Xu , Yutian Lei , Liangjun Zhang

We present ProgVLA, a compact vision-language-action (VLA) model designed for reliable robot manipulation under tight compute and memory budgets. The model specifically focuses on efficiently processing long multi-modal sequences by…

Robotics · Computer Science 2026-05-28 Seungsu Kim , Jinyoung Choi , Seungmin Baek , Jean-Michel Renders

Vision-Language-Action (VLA) models have recently shown strong potential for robot learning by following language instructions. However, in practice, language alone is often insufficient to precisely convey human intent. It is difficult to…

We propose Heterogeneous Masked Autoregression (HMA) for modeling action-video dynamics to generate high-quality data and evaluation in scaling robot learning. Building interactive video world models and policies for robotics is difficult…

Robotics · Computer Science 2025-02-07 Lirui Wang , Kevin Zhao , Chaoqi Liu , Xinlei Chen

We present a unified perspective on tackling various human-centric video tasks by learning human motion representations from large-scale and heterogeneous data resources. Specifically, we propose a pretraining stage in which a motion…

Computer Vision and Pattern Recognition · Computer Science 2023-08-15 Wentao Zhu , Xiaoxuan Ma , Zhaoyang Liu , Libin Liu , Wayne Wu , Yizhou Wang

Humanoid robots have shown success in locomotion and manipulation. Despite these basic abilities, humanoids are still required to quickly understand human instructions and react based on human interaction signals to become valuable…

Human teams are able to easily perform collaborative manipulation tasks. However, for a robot and human to simultaneously manipulate an extended object is a difficult task using existing methods from the literature. Our approach in this…

Robotics · Computer Science 2020-01-07 Erich Mielke , Eric Townsend , David Wingate , Marc D. Killpack

Building a generalist robot that can perceive, reason, and act across diverse tasks remains an open challenge, especially for dexterous manipulation. A major bottleneck lies in the scarcity of large-scale, action-annotated data for…

Robotics · Computer Science 2025-11-24 Yankai Fu , Ning Chen , Junkai Zhao , Shaozhe Shan , Guocai Yao , Pengwei Wang , Zhongyuan Wang , Shanghang Zhang

This paper presents an improved system based on our prior work, designed to create explanations for autonomous robot actions during Human-Robot Interaction (HRI). Previously, we developed a system that used Large Language Models (LLMs) to…

Egocentric videos capture how humans manipulate objects and tools, providing diverse motion cues for learning object manipulation. Unlike the costly, expert-driven manual teleoperation commonly used in training Vision-Language-Action models…

Robotics · Computer Science 2025-09-29 Tomoya Yoshida , Shuhei Kurita , Taichi Nishimura , Shinsuke Mori

Reinforcement learning (RL) holds great promise for enabling autonomous acquisition of complex robotic manipulation skills, but realizing this potential in real-world settings has been challenging. We present a human-in-the-loop…

Robotics · Computer Science 2025-03-21 Jianlan Luo , Charles Xu , Jeffrey Wu , Sergey Levine

Vision-Language-Action (VLA) models are prone to compounding errors in dexterous manipulation, where high-dimensional action spaces and contact-rich dynamics amplify small policy deviations over long horizons. While Interactive Imitation…

Robotics · Computer Science 2026-05-21 Zhuohang Li , Liqun Huang , Wei Xu , Zhengming Zhu , Nie Lin , Xiao Ma , Xinjun Sheng , Ruoshi Wen

The recent advances in instance-level detection tasks lay strong foundation for genuine comprehension of the visual scenes. However, the ability to fully comprehend a social scene is still in its preliminary stage. In this work, we focus on…

Computer Vision and Pattern Recognition · Computer Science 2019-09-24 Bingjie Xu , Junnan Li , Yongkang Wong , Mohan S. Kankanhalli , Qi Zhao

We present Vision in Action (ViA), an active perception system for bimanual robot manipulation. ViA learns task-relevant active perceptual strategies (e.g., searching, tracking, and focusing) directly from human demonstrations. On the…

Robotics · Computer Science 2025-06-19 Haoyu Xiong , Xiaomeng Xu , Jimmy Wu , Yifan Hou , Jeannette Bohg , Shuran Song

In mobile robot shared control, effectively understanding human motion intention is critical for seamless human-robot collaboration. This paper presents a novel shared control framework featuring planning-level intention prediction. A path…

Robotics · Computer Science 2025-11-13 Jinyu Zhang , Lijun Han , Feng Jian , Lingxi Zhang , Hesheng Wang

Fine-grained understanding of human actions and poses in videos is essential for human-centric AI applications. In this work, we introduce ActionArt, a fine-grained video-caption dataset designed to advance research in human-centric…

Computer Vision and Pattern Recognition · Computer Science 2025-04-28 Yi-Xing Peng , Qize Yang , Yu-Ming Tang , Shenghao Fu , Kun-Yu Lin , Xihan Wei , Wei-Shi Zheng

Representation learning approaches for robotic manipulation have boomed in recent years. Due to the scarcity of in-domain robot data, prevailing methodologies tend to leverage large-scale human video datasets to extract generalizable…

Learning actions from human demonstration is an emerging trend for designing intelligent robotic systems, which can be referred as video to command. The performance of such approach highly relies on the quality of video captioning. However,…

Computer Vision and Pattern Recognition · Computer Science 2019-09-11 Shuo Yang , Wei Zhang , Weizhi Lu , Hesheng Wang , Yibin Li

Human emotions are expressed through multiple modalities, including verbal and non-verbal information. Moreover, the affective states of human users can be the indicator for the level of engagement and successful interaction, suitable for…

Robotics · Computer Science 2021-10-12 Baijun Xie , Chung Hyuk Park
‹ Prev 1 8 9 10 Next ›