English
Related papers

Related papers: Where2Act: From Pixels to Actions for Articulated …

200 papers

Pushing is an essential non-prehensile manipulation skill used for tasks ranging from pre-grasp manipulation to scene rearrangement, reasoning about object relations in the scene, and thus pushing actions have been widely studied in…

Robotics · Computer Science 2024-05-22 Ahmet E. Tekden , Aykut Erdem , Erkut Erdem , Tamim Asfour , Emre Ugur

Deep learning requires large amounts of training data to be effective. For the task of object segmentation, manually labeling data is very expensive, and hence interactive methods are needed. Following recent approaches, we develop an…

Computer Vision and Pattern Recognition · Computer Science 2018-05-14 Sabarinath Mahadevan , Paul Voigtlaender , Bastian Leibe

From dishwashers to cabinets, humans interact with articulated objects every day, and for a robot to assist in common manipulation tasks, it must learn a representation of articulation. Recent deep learning learning methods can provide…

Robotics · Computer Science 2023-09-29 Russell Buchanan , Adrian Röfer , João Moura , Abhinav Valada , Sethu Vijayakumar

In computer vision, video-based approaches have been widely explored for the early classification and the prediction of actions or activities. However, it remains unclear whether this modality (as compared to 3D kinematics) can still be…

Computer Vision and Pattern Recognition · Computer Science 2017-08-04 Andrea Zunino , Jacopo Cavazza , Atesh Koul , Andrea Cavallo , Cristina Becchio , Vittorio Murino

Representing a scene and its constituent objects from raw sensory data is a core ability for enabling robots to interact with their environment. In this paper, we propose a novel approach for scene understanding, leveraging a hierarchical…

Robotics · Computer Science 2023-02-08 Toon Van de Maele , Tim Verbelen , Pietro Mazzaglia , Stefano Ferraro , Bart Dhoedt

Accurate localization in diverse environments is a fundamental challenge in computer vision and robotics. The task involves determining a sensor's precise position and orientation, typically a camera, within a given space. Traditional…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Luca Di Giammarino , Boyang Sun , Giorgio Grisetti , Marc Pollefeys , Hermann Blum , Daniel Barath

Although modern object detection and classification models achieve high accuracy, these are typically constrained in advance on a fixed train set and are therefore not flexible to deal with novel, unseen object categories. Moreover, these…

Artificial Intelligence · Computer Science 2021-08-27 Toon Van de Maele , Tim Verbelen , Ozan Catal , Bart Dhoedt

We study the problem of learning a generalizable action policy for an intelligent agent to actively approach an object of interest in an indoor environment solely from its visual inputs. While scene-driven or recognition-driven visual…

Robotics · Computer Science 2019-03-08 Xin Ye , Zhe Lin , Joon-Young Lee , Jianming Zhang , Shibin Zheng , Yezhou Yang

We present an open-source, real-time implementation of SemanticPaint, a system for geometric reconstruction, object-class segmentation and learning of 3D scenes. Using our system, a user can walk into a room wearing a depth camera and a…

Understanding human actions in visual data is tied to advances in complementary research areas including object recognition, human dynamics, domain adaptation and semantic segmentation. Over the last decade, human action analysis evolved…

Computer Vision and Pattern Recognition · Computer Science 2017-02-02 Samitha Herath , Mehrtash Harandi , Fatih Porikli

Intelligent agents need to select long sequences of actions to solve complex tasks. While humans easily break down tasks into subgoals and reach them through millions of muscle commands, current artificial intelligence is limited to tasks…

Artificial Intelligence · Computer Science 2022-06-10 Danijar Hafner , Kuang-Huei Lee , Ian Fischer , Pieter Abbeel

Discovering physical laws directly from high-dimensional visual data is a long-standing human pursuit but remains a formidable challenge for machines, representing a fundamental goal of scientific intelligence. This task is inherently…

Computational Engineering, Finance, and Science · Computer Science 2026-02-24 Ruikun Li , Jun Yao , Yingfan Hua , Shixiang Tang , Biqing Qi , Bin Liu , Wanli Ouyang , Yan Lu

Planning has been very successful for control tasks with known environment dynamics. To leverage planning in unknown environments, the agent needs to learn the dynamics from interactions with the world. However, learning dynamics models…

Machine Learning · Computer Science 2019-06-06 Danijar Hafner , Timothy Lillicrap , Ian Fischer , Ruben Villegas , David Ha , Honglak Lee , James Davidson

A major bottleneck for developing general reinforcement learning agents is determining rewards that will yield desirable behaviors under various circumstances. We introduce a general mechanism for automatically specifying meaningful…

Machine Learning · Computer Science 2017-11-22 Ashley D. Edwards , Charles L. Isbell

Rendering articulated objects while controlling their poses is critical to applications such as virtual reality or animation for movies. Manipulating the pose of an object, however, requires the understanding of its underlying structure,…

Computer Vision and Pattern Recognition · Computer Science 2022-04-07 Atsuhiro Noguchi , Umar Iqbal , Jonathan Tremblay , Tatsuya Harada , Orazio Gallo

Human actions often involve complex interactions across several inter-related objects in the scene. However, existing approaches to fine-grained video understanding or visual relationship detection often rely on single object representation…

Computer Vision and Pattern Recognition · Computer Science 2018-03-22 Chih-Yao Ma , Asim Kadav , Iain Melvin , Zsolt Kira , Ghassan AlRegib , Hans Peter Graf

One of the most exciting applications of vision models involve pixel-level reasoning. Despite the abundance of vision foundation models, we still lack representations that effectively embed spatio-temporal properties of visual scenes at the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Nikita Araslanov , Martin Sundermeyer , Hidenobu Matsuki , David Joseph Tan , Federico Tombari

We propose a planning and perception mechanism for a robot (agent), that can only observe the underlying environment partially, in order to solve an image classification problem. A three-layer architecture is suggested that consists of a…

Machine Learning · Computer Science 2019-09-24 Hossein K. Mousavi , Guangyi Liu , Weihang Yuan , Martin Takáč , Héctor Muñoz-Avila , Nader Motee

Latent actions serve as an intermediate representation that enables consistent modeling of vision-language-action (VLA) models across heterogeneous datasets. However, approaches to supervising VLAs with latent actions are fragmented and…

Robotics · Computer Science 2026-05-07 Yihan Lin , Haoyang Li , Yang Li , Haitao Shen , Yihan Zhao , Chao Shao , Jing Zhang

We propose a model of a learning agent whose interaction with the environment is governed by a simulation-based projection, which allows the agent to project itself into future situations before it takes real action. Projective simulation…

Adaptation and Self-Organizing Systems · Physics 2015-03-19 Hans J. Briegel , Gemma De las Cuevas