English
Related papers

Related papers: COPILOT: Human-Environment Collision Prediction an…

200 papers

In today's Human-Robot Interaction (HRI) scenarios, a prevailing tendency exists to assume that the robot shall cooperate with the closest individual or that the scene involves merely a singular human actor. However, in realistic scenarios,…

Computer Vision and Pattern Recognition · Computer Science 2024-02-01 Federico Rollo , Andrea Zunino , Nikolaos Tsagarakis , Enrico Mingo Hoffman , Arash Ajoudani

While current vision algorithms excel at many challenging tasks, it is unclear how well they understand the physical dynamics of real-world environments. Here we introduce Physion, a dataset and benchmark for rigorously evaluating the…

In egocentric action recognition a single population model is typically trained and subsequently embodied on a head-mounted device, such as an augmented reality headset. While this model remains static for new users and environments, we…

Computer Vision and Pattern Recognition · Computer Science 2023-07-13 Matthias De Lange , Hamid Eghbalzadeh , Reuben Tan , Michael Iuzzolino , Franziska Meier , Karl Ridgeway

The advancement of robot learning is currently hindered by the scarcity of large-scale, high-quality datasets. While established data collection methods such as teleoperation and universal manipulation interfaces dominate current datasets,…

Predictive models have been at the core of many robotic systems, from quadrotors to walking robots. However, it has been challenging to develop and apply such models to practical robotic manipulation due to high-dimensional sensory…

Robotics · Computer Science 2020-09-14 Lucas Manuelli , Yunzhu Li , Pete Florence , Russ Tedrake

To serve as a scalable data source for embodied AI, world models should act as true simulators that infer interaction dynamics strictly from user actions, rather than mere conditional video generators relying on privileged future object…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Dayou Li , Lulin Liu , Bangya Liu , Shijie Zhou , Jiu Feng , Ziqi Lu , Minghui Zheng , Chenyu You , Zhiwen Fan

Although First Person Vision systems can sense the environment from the user's perspective, they are generally unable to predict his intentions and goals. Since human activities can be decomposed in terms of atomic actions and interactions…

Computer Vision and Pattern Recognition · Computer Science 2019-04-11 Antonino Furnari , Sebastiano Battiato , Kristen Grauman , Giovanni Maria Farinella

Detecting small objects in video streams of head-worn augmented reality devices in near real-time is a huge challenge: training data is typically scarce, the input video stream can be of limited quality, and small objects are notoriously…

Computer Vision and Pattern Recognition · Computer Science 2021-06-14 Hooman Tavakoli , Snehal Walunj , Parsha Pahlevannejad , Christiane Plociennik , Martin Ruskowski

Predicting future human behavior from egocentric videos is a challenging but critical task for human intention understanding. Existing methods for forecasting 2D hand positions rely on visual representations and mainly focus on hand-object…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Masashi Hatano , Ryo Hachiuma , Hideo Saito

We introduce an approach for pre-training egocentric video models using large-scale third-person video datasets. Learning from purely egocentric data is limited by low dataset scale and diversity, while using purely exocentric…

Computer Vision and Pattern Recognition · Computer Science 2021-04-19 Yanghao Li , Tushar Nagarajan , Bo Xiong , Kristen Grauman

Understanding multimodal signals in egocentric vision, such as RGB video, depth, camera poses, and gaze, is essential for applications in augmented reality, robotics, and human-computer interaction, enabling systems to better interpret the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Gen Li , Yutong Chen , Yiqian Wu , Kaifeng Zhao , Marc Pollefeys , Siyu Tang

We propose the use of a proportional-derivative (PD) control based policy learned via reinforcement learning (RL) to estimate and forecast 3D human pose from egocentric videos. The method learns directly from unsegmented egocentric videos…

Computer Vision and Pattern Recognition · Computer Science 2019-08-06 Ye Yuan , Kris Kitani

Accurate 3D object detection in real-world environments requires a huge amount of annotated data with high quality. Acquiring such data is tedious and expensive, and often needs repeated effort when a new sensor is adopted or when the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Jinsu Yoo , Zhenyang Feng , Tai-Yu Pan , Yihong Sun , Cheng Perng Phoo , Xiangyu Chen , Mark Campbell , Kilian Q. Weinberger , Bharath Hariharan , Wei-Lun Chao

Different video understanding tasks are typically treated in isolation, and even with distinct types of curated data (e.g., classifying sports in one dataset, tracking animals in another). However, in wearable cameras, the immersive…

Computer Vision and Pattern Recognition · Computer Science 2023-04-10 Zihui Xue , Yale Song , Kristen Grauman , Lorenzo Torresani

Human comprehension of a video stream is naturally broad: in a few instants, we are able to understand what is happening, the relevance and relationship of objects, and forecast what will follow in the near future, everything all at once.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Simone Alberto Peirone , Francesca Pistilli , Antonio Alliegro , Giuseppe Averta

We propose a self-supervised algorithm to learn representations from egocentric video data. Recently, significant efforts have been made to capture humans interacting with their own environments as they go about their daily activities. In…

Computer Vision and Pattern Recognition · Computer Science 2022-09-28 Himangi Mittal , Pedro Morgado , Unnat Jain , Abhinav Gupta

Many works in collaborative robotics and human-robot interaction focuses on identifying and predicting human behaviour while considering the information about the robot itself as given. This can be the case when sensors and the robot are…

Robots operating in populated environments encounter many different types of people, some of whom might have an advanced need for cautious interaction, because of physical impairments or their advanced age. Robots therefore need to…

Robotics · Computer Science 2017-08-03 Andres Vasquez , Marina Kollmitz , Andreas Eitel , Wolfram Burgard

Sampling-based motion planning is an effective tool to compute safe trajectories for automated vehicles in complex environments. However, a fast convergence to the optimal solution can only be ensured with the use of problem-specific…

Robotics · Computer Science 2019-02-04 Holger Banzhaf , Paul Sanzenbacher , Ulrich Baumann , J. Marius Zöllner

Walking has always been a primary mode of transportation and is recognized as an essential activity for maintaining good health. Despite the need for safe walking conditions in urban environments, sidewalks are frequently obstructed by…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Marios Thoma , Zenonas Theodosiou , Harris Partaourides , Vassilis Vassiliades , Loizos Michael , Andreas Lanitis