English
Related papers

Related papers: ViPRA: Video Prediction for Robot Actions

200 papers

Visual-textual understanding is essential for language-guided robot manipulation. Recent works leverage pre-trained vision-language models to measure the similarity between encoded visual observations and textual instructions, and then…

Robotics · Computer Science 2025-09-30 Chaoran Zhu , Hengyi Wang , Yik Lung Pang , Changjae Oh

Vision-language-action (VLA) models can enable broad open world generalization, but require large and diverse datasets. It is appealing to consider whether some of this data can come from human videos, which cover diverse real-world…

Human videos contain rich manipulation priors, but using them for robot learning remains difficult because raw observations entangle scene understanding, human motion, and embodiment-specific action. We introduce MoT-HRA, a hierarchical…

Robotics · Computer Science 2026-05-22 Yifan Xie , YuAn Wang , Guangyu Chen , Jinkun Liu , Yu Sun , Wenbo Ding

One of the great promises of robot learning systems is that they will be able to learn from their mistakes and continuously adapt to ever-changing environments. Despite this potential, most of the robot learning systems today are deployed…

Machine Learning · Computer Science 2020-08-03 Ryan Julian , Benjamin Swanson , Gaurav S. Sukhatme , Sergey Levine , Chelsea Finn , Karol Hausman

For robots to operate effectively and safely alongside humans, they must be able to understand the progress of ongoing actions. This ability, known as action progress prediction, is critical for tasks ranging from timely assistance to…

Robotics · Computer Science 2026-03-03 Elena Zoppellari , Federico Becattini , Marco Fiorucci , Lamberto Ballan

Many Vision-Language-Action (VLA) models flatten image patches into a 1D token sequence, weakening the 2D spatial cues needed for precise manipulation. We introduce IVRA, a lightweight, training-free method that improves spatial…

Robotics · Computer Science 2026-01-23 Jongwoo Park , Kanchana Ranasinghe , Jinhyeok Jang , Cristina Mata , Yoo Sung Jang , Michael S Ryoo

Vision-Language-Action (VLA) policies are typically deployed with asynchronous inference: the robot executes a previously predicted action chunk while the model computes the next one. This creates a prediction-execution misalignment: the…

Robotics · Computer Science 2026-05-20 Yixiang Zhu , Yonghao Chen , Rui Meng , Jingyu Guo , Jiaxiang Zou , Zijie Yang , Taowen Wang , Xinyu Chen

Trustworthy robot behavior requires not only high levels of task success but also that the robot can reliably quantify how likely it is to succeed. To this end, we present a first-of-its-kind study of confidence calibration in…

Robotics · Computer Science 2025-12-23 Thomas P Zollo , Richard Zemel

Learning from demonstration is a powerful method for teaching robots new skills, and having more demonstration data often improves policy learning. However, the high cost of collecting demonstration data is a significant bottleneck. Videos,…

Robotics · Computer Science 2024-07-15 Chuan Wen , Xingyu Lin , John So , Kai Chen , Qi Dou , Yang Gao , Pieter Abbeel

Real robot data collection for imitation learning has led to significant advancements in robotic manipulation. However, the requirement for robot hardware in the process fundamentally constrains the scale of the data. In this paper, we…

Despite progress, Vision-Language-Action models (VLAs) are limited by a scarcity of large-scale, diverse robot data. While human manipulation videos offer a rich alternative, existing methods are forced to choose between small,…

Robotics · Computer Science 2026-02-26 Hao Luo , Ye Wang , Wanpeng Zhang , Haoqi Yuan , Yicheng Feng , Haiweng Xu , Sipeng Zheng , Zongqing Lu

Most existing vision-language-action (VLA) models for robotic manipulation lack progress awareness, typically relying on hand-crafted heuristics for task termination. This limitation is particularly severe in long-horizon tasks involving…

Robotics · Computer Science 2026-03-31 Hongyu Yan , Qiwei Li , Jiaolong Yang , Yadong Mu

Building generalist robot policies that can handle diverse tasks in open-ended environments is a central challenge in robotics. To leverage knowledge from large-scale pretraining, prior work (VLA) has typically built generalist policies…

Robotics · Computer Science 2026-05-14 Jianke Zhang , Yucheng Hu , Yanjiang Guo , Xiaoyu Chen , Yichen Liu , Wenna Chen , Chaochao Lu , Jianyu Chen

The pre-training of visual representations has enhanced the efficiency of robot learning. Due to the lack of large-scale in-domain robotic datasets, prior works utilize in-the-wild human videos to pre-train robotic visual representation.…

Robotics · Computer Science 2024-10-31 Guangqi Jiang , Yifei Sun , Tao Huang , Huanyu Li , Yongyuan Liang , Huazhe Xu

Robots in shared workspaces must interpret human actions from partial, ambiguous observations, where overconfident early predictions can lead to unsafe or disruptive interaction. This challenge is amplified in egocentric views, where…

Robotics · Computer Science 2026-03-13 Zhaoda Du , Michael Bowman , Qiaojie Zheng , Xiaoli Zhang

With the advancement in computer vision deep learning, systems now are able to analyze an unprecedented amount of rich visual information from videos to enable applications such as autonomous driving, socially-aware robot assistant and…

Computer Vision and Pattern Recognition · Computer Science 2021-07-19 Junwei Liang

Vision-language-action (VLA) models finetuned from vision-language models (VLMs) hold the promise of leveraging rich pretrained representations to build generalist robots across diverse tasks and environments. However, direct fine-tuning on…

Robotics · Computer Science 2025-09-18 Shresth Grover , Akshay Gopalkrishnan , Bo Ai , Henrik I. Christensen , Hao Su , Xuanlin Li

Understanding human perceptions of robot performance is crucial for designing socially intelligent robots that can adapt to human expectations. Current approaches often rely on surveys, which can disrupt ongoing human-robot interactions. As…

Recent unsupervised pre-training methods have shown to be effective on language and vision domains by learning useful representations for multiple downstream tasks. In this paper, we investigate if such unsupervised pre-training methods can…

Computer Vision and Pattern Recognition · Computer Science 2022-06-20 Younggyo Seo , Kimin Lee , Stephen James , Pieter Abbeel

Vision Language Action (VLA) models derive their generalization capability from diverse training data, yet collecting embodied robot interaction data remains prohibitively expensive. In contrast, human demonstration videos are far more…