中文
相关论文

相关论文: Vidarc: Embodied Video Diffusion Model for Closed-…

200 篇论文

Learning a generalizable bimanual manipulation policy is extremely challenging for embodied agents due to the large action space and the need for coordinated arm movements. Existing approaches rely on Vision-Language-Action (VLA) models to…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Chenyou Fan , Fangzheng Yan , Chenjia Bai , Jiepeng Wang , Chi Zhang , Zhen Wang , Xuelong Li

Embodied world models aim to predict and interact with the physical world through visual observations and actions. However, existing models struggle to accurately translate low-level actions (e.g., joint positions) into precise robotic…

机器人学 · 计算机科学 2026-04-01 Taiyi Su , Jian Zhu , Yaxuan Li , Chong Ma , Jianjun Zhang , Zitai Huang , Hanli Wang , Yi Xu

Despite significant progress in robotics and embodied AI in recent years, deploying robots for long-horizon tasks remains a great challenge. Majority of prior arts adhere to an open-loop philosophy and lack real-time feedback, leading to…

机器人学 · 计算机科学 2024-10-17 Qingwen Bu , Jia Zeng , Li Chen , Yanchao Yang , Guyue Zhou , Junchi Yan , Ping Luo , Heming Cui , Yi Ma , Hongyang Li

Predicting and anticipating future outcomes or reasoning about missing information in a sequence are critical skills for agents to be able to make intelligent decisions. This requires strong, temporally coherent generative capabilities.…

计算机视觉与模式识别 · 计算机科学 2022-11-15 Tobias Höppe , Arash Mehrjou , Stefan Bauer , Didrik Nielsen , Andrea Dittadi

Recent advancements in video generation have seen a shift towards unified, transformer-based foundation models that can handle multiple conditional inputs in-context. However, these models have primarily focused on modalities like text,…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Wenze Liu , Weicai Ye , Minghong Cai , Quande Liu , Xintao Wang , Xiangyu Yue

Human videos are a scalable source of training data for robot learning. However, humans and robots significantly differ in embodiment, making many human actions infeasible for direct execution on a robot. Still, these demonstrations convey…

Imitation learning is a powerful paradigm for robot skill acquisition, yet conventional demonstration methods--such as kinesthetic teaching and teleoperation--are cumbersome, hardware-heavy, and disruptive to workflows. Recently, passive…

机器人学 · 计算机科学 2025-09-30 Rohan Walia , Yusheng Wang , Ralf Römer , Masahiro Nishio , Angela P. Schoellig , Jun Ota

A key challenge with procedure planning in instructional videos lies in how to handle a large decision space consisting of a multitude of action types that belong to various tasks. To understand real-world video content, an AI agent must…

计算机视觉与模式识别 · 计算机科学 2023-09-15 Fen Fang , Yun Liu , Ali Koksal , Qianli Xu , Joo-Hwee Lim

Learning to navigate in unstructured environments is a challenging task for robots. While reinforcement learning can be effective, it often requires extensive data collection and can pose risk. Learning from expert demonstrations, on the…

机器人学 · 计算机科学 2024-12-31 Nimrod Curtis , Osher Azulay , Avishai Sintov

Lately, there has been growing interest in adapting vision-language models (VLMs) to image and third-person video classification due to their success in zero-shot recognition. However, the adaptation of these models to egocentric videos has…

计算机视觉与模式识别 · 计算机科学 2024-04-01 Anna Kukleva , Fadime Sener , Edoardo Remelli , Bugra Tekin , Eric Sauser , Bernt Schiele , Shugao Ma

A novel concept of vision-based intelligent control of robotic arms is developed here in this work. This work enables the controlling of robotic arms motion only with visual inputs, that is, controlling by showing the videos of correct…

计算机视觉与模式识别 · 计算机科学 2021-01-28 Debarati B. Chakraborty , Mukesh Sharma , Bhaskar Vijay

Embodied control requires agents to leverage multi-modal pre-training to quickly learn how to act in new environments, where video demonstrations contain visual and motion details needed for low-level perception and control, and language…

机器学习 · 计算机科学 2023-04-20 Yao Mu , Shunyu Yao , Mingyu Ding , Ping Luo , Chuang Gan

Pre-training for Reinforcement Learning (RL) with purely video data is a valuable yet challenging problem. Although in-the-wild videos are readily available and inhere a vast amount of prior world knowledge, the absence of action…

计算机视觉与模式识别 · 计算机科学 2024-11-06 Hao Luo , Bohan Zhou , Zongqing Lu

Discrete diffusion models have recently shown great promise for modeling complex discrete data, with masked diffusion models (MDMs) offering a compelling trade-off between quality and generation speed. MDMs denoise by progressively…

机器学习 · 计算机科学 2026-04-15 Tianyu Xie , Shuchen Xue , Zijin Feng , Tianyang Hu , Jiacheng Sun , Zhenguo Li , Cheng Zhang

We introduce UMI-on-Air, a framework for embodiment-aware deployment of embodiment-agnostic manipulation policies. Our approach leverages diverse, unconstrained human demonstrations collected with a handheld gripper (UMI) to train…

机器人学 · 计算机科学 2026-03-17 Harsh Gupta , Xiaofeng Guo , Huy Ha , Chuer Pan , Muqing Cao , Dongjae Lee , Sebastian Scherer , Shuran Song , Guanya Shi

Learning robot manipulation from abundant human videos offers a scalable alternative to costly robot-specific data collection. However, domain gaps across visual, morphological, and physical aspects hinder direct imitation. To effectively…

机器人学 · 计算机科学 2025-09-16 Yangcen Liu , Woo Chul Shin , Yunhai Han , Zhenyang Chen , Harish Ravichandar , Danfei Xu

Visuomotor imitation learning policies enable robots to efficiently acquire manipulation skills from visual demonstrations. However, as scene complexity and visual distractions increase, policies that perform well in simple settings often…

World models, which predict future transitions from past observation and action sequences, have shown great promise for improving data efficiency in sequential decision-making. However, existing world models often require extensive…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Siqiao Huang , Jialong Wu , Qixing Zhou , Shangchen Miao , Mingsheng Long

Robot manipulation research still suffers from significant data scarcity: even the largest robot datasets are orders of magnitude smaller and less diverse than those that fueled recent breakthroughs in language and vision. We introduce…

机器人学 · 计算机科学 2026-05-29 Marion Lepert , Jiaying Fang , Jeannette Bohg

We present a method to generate video-action pairs that follow text instructions, starting from an initial image observation and the robot's joint states. Our approach automatically provides action labels for video diffusion models,…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Liudi Yang , Yang Bai , George Eskandar , Fengyi Shen , Mohammad Altillawi , Dong Chen , Ziyuan Liu , Abhinav Valada