English
Related papers

Related papers: Long-Horizon Manipulation via Trace-Conditioned VL…

200 papers

This work addresses the problem of long-horizon task planning with the Large Language Model (LLM) in an open-world household environment. Existing works fail to explicitly track key objects and attributes, leading to erroneous decisions in…

Robotics · Computer Science 2024-04-23 Siwei Chen , Anxing Xiao , David Hsu

As aerial platforms evolve from passive observers to active manipulators, the challenge shifts toward designing intuitive interfaces that allow non-expert users to command these systems naturally. This work introduces a novel concept of…

In this paper we propose a new framework - MoViLan (Modular Vision and Language) for execution of visually grounded natural language instructions for day to day indoor household tasks. While several data-driven, end-to-end learning…

Computer Vision and Pattern Recognition · Computer Science 2021-01-21 Homagni Saha , Fateme Fotouhif , Qisai Liu , Soumik Sarkar

We study rigid-body motion planning through multiple sequential narrow openings, which requires long-horizon geometric reasoning because the configuration used to traverse an early opening constrains the set of reachable configurations for…

Robotics · Computer Science 2026-03-18 Al Jaber Mahmud , Xuan Wang

The ability of Language Models (LMs) to understand natural language makes them a powerful tool for parsing human instructions into task plans for autonomous robots. Unlike traditional planning methods that rely on domain-specific knowledge…

Achieving human-like dexterous manipulation remains a major challenge for general-purpose robots. While Vision-Language-Action (VLA) models show potential in learning skills from demonstrations, their scalability is limited by scarce…

Robotics · Computer Science 2025-12-16 Yu Cui , Yujian Zhang , Lina Tao , Yang Li , Xinyu Yi , Zhibin Li

Vision-Language-Action (VLA) models aim to control robots for manipulation from visual observations and natural-language instructions. However, existing hierarchical and autoregressive paradigms often introduce architectural overhead,…

Understanding how Vision-Language-Action (VLA) models transform multimodal knowledge into embodied control remains an open challenge. We present VLA-Trace, a progressive diagnostic framework that analyzes VLA models through a unified…

Artificial Intelligence · Computer Science 2026-05-29 Haoyuan Shi , Xiancong Ren , Yingji Zhang , Qinfan Zhang , Jiayu Hu , Haozhe Shan , Han Dong , Jinpeng Lu , Yinda Chen , Yi Zhang , Yong Dai , Xiaozhu Ju

Precision-critical manipulation requires both global trajectory organization and local execution correction, yet most vision-language-action (VLA) policies generate actions within a single unified space. This monolithic formulation forces…

Robotics · Computer Science 2026-04-21 Tingzheng Jia , Kan Guo , Lanping Qian , Yongli Hu , Daxin Tian , Guixian Qu , Chunmian Lin , Baocai Yin , Jiapu Wang

Robotic manipulation policies often degrade over extended horizons, yet existing benchmarks provide limited insight into why such failures occur. Most prior benchmarks are either simulation-based or report aggregate success, making it…

Robotics · Computer Science 2026-04-21 Xueyao Chen , Jingkai Jia , Tong Yang , Yibo Fu , Wei Li , Wenqiang Zhang

Large language model (LLM) agents have recently demonstrated strong capabilities in interactive decision-making, yet they remain fundamentally limited in long-horizon tasks that require structured planning and reliable execution. Existing…

Artificial Intelligence · Computer Science 2026-05-06 Hongbo Jin , Rongpeng Zhu , Jiayu Ding , Guibo Luo , Ge Li

Most existing vision-language-action (VLA) models for robotic manipulation lack progress awareness, typically relying on hand-crafted heuristics for task termination. This limitation is particularly severe in long-horizon tasks involving…

Robotics · Computer Science 2026-03-31 Hongyu Yan , Qiwei Li , Jiaolong Yang , Yadong Mu

Vision-Language-Action (VLA) models are promising for generalist robot manipulation but remain brittle in out-of-distribution (OOD) settings, especially with limited real-robot data. To resolve the generalization bottleneck, we introduce a…

Existing Vision-Language models (VLMs) estimate either long-term trajectory waypoints or a set of control actions as a reactive solution for closed-loop planning based on their rich scene comprehension. However, these estimations are coarse…

Robotics · Computer Science 2024-04-01 Pranjal Paul , Anant Garg , Tushar Choudhary , Arun Kumar Singh , K. Madhava Krishna

Assisting humans in open-world outdoor environments requires robots to translate high-level natural-language intentions into safe, long-horizon, and socially compliant navigation behavior. Existing map-based methods rely on costly pre-built…

Vision-Language-Action (VLA) models achieve over 95% success on standard benchmarks. However, through systematic experiments, we find that current state-of-the-art VLA models largely ignore language instructions. Prior work lacks: (1)…

Robotics · Computer Science 2026-03-03 Yuchen Hou , Lin Zhao

Real-world robotic manipulation tasks remain an elusive challenge, since they involve both fine-grained environment interaction, as well as the ability to plan for long-horizon goals. Although deep reinforcement learning (RL) methods have…

Machine Learning · Computer Science 2023-03-20 Núria Armengol Urpí , Marco Bagatella , Otmar Hilliges , Georg Martius , Stelian Coros

Vision-Language-Action (VLA) models are prone to compounding errors in dexterous manipulation, where high-dimensional action spaces and contact-rich dynamics amplify small policy deviations over long horizons. While Interactive Imitation…

Robotics · Computer Science 2026-05-21 Zhuohang Li , Liqun Huang , Wei Xu , Zhengming Zhu , Nie Lin , Xiao Ma , Xinjun Sheng , Ruoshi Wen

Vision-Language-Action (VLA) models inherit strong priors from pretrained Vision-Language Models (VLMs), but naive fine-tuning often disrupts these representations and harms generalization. Existing fixes -- freezing modules or applying…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Chengyue Huang , Mellon M. Zhang , Robert Azarcon , Glen Chou , Zsolt Kira

Vision-Language-Action (VLA) models are increasingly expected to not only complete robot tasks, but also follow human instructions about how those tasks should be executed. However, existing robot datasets usually pair trajectories with…

‹ Prev 1 4 5 6 7 8 10 Next ›