中文
相关论文

相关论文: Instruct2Act: From Human Instruction to Actions Se…

200 篇论文

We present a novel framework, Action Progression Network (APN), for temporal action detection (TAD) in videos. The framework locates actions in videos by detecting the action evolution process. To encode the action evolution, we quantify a…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Chongkai Lu , Man-Wai Mak , Ruimin Li , Zheru Chi , Hong Fu

Real-life robot navigation involves more than just reaching a destination; it requires optimizing movements while addressing scenario-specific goals. An intuitive way for humans to express these goals is through abstract cues like verbal…

A fundamental requirement for real-world robotic deployment is the ability to understand and respond to natural language instructions. Existing language-conditioned manipulation tasks typically assume that instructions are perfectly aligned…

Language agents have demonstrated autonomous decision-making abilities by reasoning with foundation models. Recently, efforts have been made to train language agents for performance improvement, with multi-step reasoning and action…

人工智能 · 计算机科学 2024-04-02 Zonghan Yang , Peng Li , Ming Yan , Ji Zhang , Fei Huang , Yang Liu

Human-centred systems require an understanding of human actions in the physical world. Temporally extended sequences of actions are intentional and structured, yet existing methods for recognising what actions are performed often do not…

人工智能 · 计算机科学 2026-04-21 Rimvydas Rubavicius , Manisha Dubey , N. Siddharth , Subramanian Ramamoorthy

A longstanding goal of artificial intelligence is to create artificial agents capable of learning to perform tasks that require sequential decision making. Importantly, while it is the artificial agent that learns and acts, it is still up…

人工智能 · 计算机科学 2021-07-14 Ruohan Zhang , Faraz Torabi , Garrett Warnell , Peter Stone

Collecting diverse, high-quality manipulation data for Vision-Language-Action (VLA) model training remains prohibitively expensive for many research groups, as existing teleoperation frameworks rely on specialized hardware or are tightly…

机器人学 · 计算机科学 2026-05-05 Om Mandhane , Bipin Yadav , Sangeetha Prasanna Ram , Gopalakrishnan Narayanan

Understanding human actions in visual data is tied to advances in complementary research areas including object recognition, human dynamics, domain adaptation and semantic segmentation. Over the last decade, human action analysis evolved…

计算机视觉与模式识别 · 计算机科学 2017-02-02 Samitha Herath , Mehrtash Harandi , Fatih Porikli

Understanding human perceptions of robot performance is crucial for designing socially intelligent robots that can adapt to human expectations. Current approaches often rely on surveys, which can disrupt ongoing human-robot interactions. As…

机器人学 · 计算机科学 2025-12-03 Qiping Zhang , Nathan Tsoi , Mofeed Nagib , Booyeon Choi , Jie Tan , Hao-Tien Lewis Chiang , Marynel Vázquez

We introduce Act2Vec, a general framework for learning context-based action representation for Reinforcement Learning. Representing actions in a vector space help reinforcement learning algorithms achieve better performance by grouping…

人工智能 · 计算机科学 2019-05-21 Guy Tennenholtz , Shie Mannor

Vision-language-action (VLA) reasoning tasks require agents to interpret multimodal instructions, perform long-horizon planning, and act adaptively in dynamic environments. Existing approaches typically train VLA models in an end-to-end…

计算机视觉与模式识别 · 计算机科学 2025-09-19 Chi-Pin Huang , Yueh-Hua Wu , Min-Hung Chen , Yu-Chiang Frank Wang , Fu-En Yang

Accurately manipulating articulated objects is a challenging yet important task for real robot applications. In this paper, we present a novel framework called Sim2Real$^2$ to enable the robot to manipulate an unseen articulated object to…

机器人学 · 计算机科学 2023-02-22 Liqian Ma , Jiaojiao Meng , Shuntao Liu , Weihang Chen , Jing Xu , Rui Chen

Behavioral skills or policies for autonomous agents are conventionally learned from reward functions, via reinforcement learning, or from demonstrations, via imitation learning. However, both modes of task specification have their…

Deployed machine learning systems face distribution drift, yet most monitoring pipelines stop at alarms and leave the response underspecified under labeling, compute, and latency constraints. We introduce Drift2Act, a drift-to-action…

机器学习 · 计算机科学 2026-03-10 Ismail Lamaakal , Chaymae Yahyati , Khalid El Makkaoui , Ibrahim Ouahbi , Yassine Maleh

The potential for agents, whether embodied or software, to learn by observing other agents performing procedures involving objects and actions is rich. Current research on automatic procedure learning heavily relies on action labels or…

计算机视觉与模式识别 · 计算机科学 2017-11-23 Luowei Zhou , Chenliang Xu , Jason J. Corso

The ability to infer pre- and postconditions of an action is vital for comprehending complex instructions, and is essential for applications such as autonomous instruction-guided agents and assistive AI that supports humans to perform…

计算与语言 · 计算机科学 2024-07-04 Te-Lin Wu , Caiqi Zhang , Qingyuan Hu , Alex Spangher , Nanyun Peng

World Action Models (WAMs) have recently emerged as a promising paradigm for robotic manipulation by jointly predicting future visual observations and future actions. However, current WAMs typically execute a fixed number of predicted…

机器人学 · 计算机科学 2026-05-12 Rui Wang , Yue Zhang , Jiehong Lin , Kuncheng Luo , Jianan Wang , Zhongrui Wang , Xiaojuan Qi

Language provides a way to break down complex concepts into digestible pieces. Recent works in robot imitation learning use language-conditioned policies that predict actions given visual observations and the high-level task specified in…

Automatic human action recognition is indispensable for almost artificial intelligent systems such as video surveillance, human-computer interfaces, video retrieval, etc. Despite a lot of progress, recognizing actions in an unknown video is…

计算机视觉与模式识别 · 计算机科学 2018-03-22 Huy-Hieu Pham , Louahdi Khoudour , Alain Crouzil , Pablo Zegers , Sergio A. Velastin