English
Related papers

Related papers: HumanVLA: Towards Vision-Language Directed Object …

200 papers

Vision-Language-Action (VLA) models have recently emerged as a powerful paradigm for robotic manipulation. Despite substantial progress enabled by large-scale pretraining and supervised fine-tuning (SFT), these models face two fundamental…

Vision Language Models (VLMs) have recently been leveraged to generate robotic actions, forming Vision-Language-Action (VLA) models. However, directly adapting a pretrained VLM for robotic control remains challenging, particularly when…

Vision-Language-Action (VLA) models have shown strong performance on embodied manipulation, yet they remain brittle under visual observation changes, paraphrased language instructions, and compounded perturbations. This limitation suggests…

Robotics · Computer Science 2026-05-20 Jingzhou Luo , Yifan Wen , Yongjie Bai , Xinshuai Song , Yang Liu , Liang Lin

Vision-language-action(VLA) models have shown great promise as generalist policies for a large range of relatively simple tasks. However, they demonstrate limited performance on more complex tasks, such as those requiring complex spatial or…

Language-driven object navigation requires agents to interpret natural language descriptions of target objects, which combine intrinsic and extrinsic attributes for instance recognition and commonsense navigation. Existing methods either…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Francesco Taioli , Shiping Yang , Sonia Raychaudhuri , Marco Cristani , Unnat Jain , Angel X Chang

Humans achieve complex manipulation through coordinated whole-body control, whereas most Vision-Language-Action (VLA) models treat robot body parts largely independently, making high-DoF humanoid control challenging and often unstable. We…

One of the fundamental quests of AI is to produce agents that coordinate well with humans. This problem is challenging, especially in domains that lack high quality human behavioral data, because multi-agent reinforcement learning (RL)…

Artificial Intelligence · Computer Science 2023-06-13 Hengyuan Hu , Dorsa Sadigh

The rise of foundation models paves the way for generalist robot policies in the physical world. Existing methods relying on text-only instructions often struggle to generalize to unseen scenarios. We argue that interleaved image-text…

Vision-Language-Action (VLA) models have emerged as a powerful paradigm for robotic manipulation. However, existing post-training methods face a dilemma between stability and exploration: Supervised Fine-Tuning (SFT) is constrained by…

Robotics · Computer Science 2026-03-17 Jiashun Li , Xiaoyu Shi , Hong Xie , Mingsheng Shang , Yun Lu

Enabling humanoid robots to reliably execute complex multi-step manipulation tasks is crucial for their effective deployment in industrial and household environments. This paper presents a hierarchical planning and control framework…

Robotics · Computer Science 2025-07-11 André Schakkal , Ben Zandonati , Zhutian Yang , Navid Azizan

Humanoid robots must adapt their contact behavior to diverse objects and tasks, yet most controllers rely on fixed, hand-tuned impedance gains and gripper settings. This paper introduces HumanoidVLM, a vision-language driven retrieval…

Robotics · Computer Science 2026-01-22 Yara Mahmoud , Yasheerah Yaqoot , Miguel Altamirano Cabrera , Dzmitry Tsetserukou

Vision-Language-Action (VLA) models are formulated to ground instructions in visual context and generate action sequences for robotic manipulation. Despite recent progress, VLA models still face challenges in learning related and reusable…

Robotics · Computer Science 2026-03-11 Ziyue Zhu , Shangyang Wu , Shuai Zhao , Zhiqiu Zhao , Shengjie Li , Yi Wang , Fang Li , Haoran Luo

This paper proposes to solve the problem of Vision-and-Language Navigation with legged robots, which not only provides a flexible way for humans to command but also allows the robot to navigate through more challenging and cluttered scenes.…

Modern Vision--Language--Action models often suffer from critical instruction-following failures in high-density manipulation environments, where task-irrelevant visual clutter dilutes attention, corrupts grounding, and substantially…

Robotics · Computer Science 2026-03-10 Zhen Liu , Xinyu Ning , Zhe Hu , XinXin Xie , Yitong Liu , Zhongzhu Pu

Human-humanoid collaboration shows significant promise for applications in healthcare, domestic assistance, and manufacturing. While compliant robot-human collaboration has been extensively developed for robotic arms, enabling compliant…

Robotics · Computer Science 2025-10-17 Yushi Du , Yixuan Li , Baoxiong Jia , Yutang Lin , Pei Zhou , Wei Liang , Yanchao Yang , Siyuan Huang

Enabling robust whole-body humanoid-object interaction (HOI) remains challenging due to motion data scarcity and the contact-rich nature. We present HDMI (HumanoiD iMitation for Interaction), a simple and general framework that learns…

Robotics · Computer Science 2025-09-30 Haoyang Weng , Yitang Li , Nikhil Sobanbabu , Zihan Wang , Zhengyi Luo , Tairan He , Deva Ramanan , Guanya Shi

Physically rearranging objects is an important capability for embodied agents. Visual room rearrangement evaluates an agent's ability to rearrange objects in a room to a desired goal based solely on visual input. We propose a simple yet…

Computer Vision and Pattern Recognition · Computer Science 2022-08-11 Brandon Trabucco , Gunnar Sigurdsson , Robinson Piramuthu , Gaurav S. Sukhatme , Ruslan Salakhutdinov

While Vision-Language-Action (VLA) models have demonstrated remarkable success in robotic manipulation, their application has largely been confined to low-degree-of-freedom end-effectors performing simple, vision-guided pick-and-place…

Robotics · Computer Science 2026-03-10 Tutian Tang , Xingyu Ji , Wanli Xing , Ce Hao , Wenqiang Xu , Lin Shao , Cewu Lu , Qiaojun Yu , Jiangmiao Pang , Kaifeng Zhang

Recent vision-language-action (VLA) models build upon vision-language foundations, and have achieved promising results and exhibit the possibility of task generalization in robot manipulation. However, due to the heterogeneity of tactile…

Robotics · Computer Science 2025-08-25 Zhengxue Cheng , Yiqian Zhang , Wenkang Zhang , Haoyu Li , Keyu Wang , Li Song , Hengdi Zhang

Vision-Language Navigation (VLN) aims to enable agents to navigate to a target location based on language instructions. Traditional VLN often follows a close-set assumption, i.e., training and test data share the same style of the input…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Yang Li , Aming Wu , Zihao Zhang , Yahong Han