English
Related papers

Related papers: EgoPush: Learning End-to-End Egocentric Multi-Obje…

200 papers

Humans constantly reason about 3D proximity, the relations between their body and surrounding objects, to guide perception and action in daily life. Whether multimodal large language models (MLLMs) can perform such embodied 3D reasoning…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Jinzhao Li , Yinuo Chen , Dongxu Piao , Panwang Pan , Yifan Yu , Dong Wang , Honglei Yan , Liang Yue , Shaofei Wang , Yixin Chen , Siyuan Huang , Miao Liu

Humans can perform previously unexperienced interactions with novel objects simply by observing others engage with them. Weakly-supervised affordance grounding mimics this process by learning to locate object regions that enable actions on…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Jiajin Tang , Zhengxuan Wei , Ge Zheng , Sibei Yang

Nonprehensile actions such as pushing are crucial for addressing multi-object rearrangement problems. Many traditional methods generate robot-centric actions, which differ from intuitive human strategies and are typically inefficient. To…

Robotics · Computer Science 2025-11-03 Kejia Ren , Gaotian Wang , Andrew S. Morgan , Lydia E. Kavraki , Kaiyu Hang

An ideal digital telepresence experience requires accurate replication of a person's body, clothing, and movements. To capture and transfer these movements into virtual reality, the egocentric (first-person) perspective can be adopted,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 G. Kutay Türkoglu , Julian Tanke , Iheb Belgacem , Lev Markhasin

Affordance, defined as the potential actions that an object offers, is crucial for embodied AI agents. For example, such knowledge directs an agent to grasp a knife by the handle for cutting or by the blade for safe handover. While existing…

Can conversational videos captured from multiple egocentric viewpoints reveal the map of a scene in a cost-efficient way? We seek to answer this question by proposing a new problem: efficiently building the map of a previously unseen 3D…

Computer Vision and Pattern Recognition · Computer Science 2023-04-24 Sagnik Majumder , Hao Jiang , Pierre Moulon , Ethan Henderson , Paul Calamia , Kristen Grauman , Vamsi Krishna Ithapu

Nowadays robots play an increasingly important role in our daily life. In human-centered environments, robots often encounter piles of objects, packed items, or isolated objects. Therefore, a robot must be able to grasp and manipulate…

Robotics · Computer Science 2022-10-06 Hamidreza Kasaei , Mohammadreza Kasaei

We present a new end-to-end learning framework to obtain detailed and spatially coherent reconstructions of multiple people from a single image. Existing multi-person methods suffer from two main drawbacks: they are often model-based and…

Computer Vision and Pattern Recognition · Computer Science 2021-04-20 Armin Mustafa , Akin Caliskan , Lourdes Agapito , Adrian Hilton

End-to-end learning is emerging as a powerful paradigm for robotic manipulation, but its effectiveness is limited by data scarcity and the heterogeneity of action spaces across robot embodiments. In particular, diverse action spaces across…

Robotics · Computer Science 2026-03-23 Erik Bauer , Elvis Nava , Robert K. Katzschmann

Autonomous vehicles and robots need to operate over a wide variety of scenarios in order to complete tasks efficiently and safely. Multi-camera self-supervised monocular depth estimation from videos is a promising way to reason about the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Takayuki Kanai , Igor Vasiljevic , Vitor Guizilini , Adrien Gaidon , Rares Ambrus

Recent advancements in learning from human demonstration have shown promising results in addressing the scalability and high cost of data collection required to train robust visuomotor policies. However, existing approaches are often…

Robotics · Computer Science 2026-04-14 Harry Freeman , Chung Hee Kim , George Kantor

Current robotic planning methods often rely on predicting multi-frame images with full pixel details. While this fine-grained approach can serve as a generic world model, it introduces two significant challenges for downstream policy…

Understanding dynamic 4D scenes from an egocentric perspective-modeling changes in 3D spatial structure over time-is crucial for human-machine interaction, autonomous navigation, and embodied intelligence. While existing egocentric datasets…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Junsheng Huang , Shengyu Hao , Bocheng Hu , Hongwei Wang , Gaoang Wang

Object manipulation for rearrangement into a specific goal state is a significant task for collaborative robots. Accurately determining object placement is a key challenge, as misalignment can increase task complexity and the risk of…

Robotics · Computer Science 2025-03-06 Guanqun Cao , Ryan Mckenna , Erich Graf , John Oyekan

Finding the camera pose is an important step in many egocentric video applications. It has been widely reported that, state of the art SLAM algorithms fail on egocentric videos. In this paper, we propose a robust method for camera pose…

Computer Vision and Pattern Recognition · Computer Science 2018-11-27 Suvam Patra , Himanshu Aggarwal , Himani Arora , Chetan Arora , Subhashis Banerjee

In this paper, we investigate the problem of anticipating future dynamics, particularly the future location of other vehicles and pedestrians, in the view of a moving vehicle. We approach two fundamental challenges: (1) the partial…

Computer Vision and Pattern Recognition · Computer Science 2020-06-09 Osama Makansi , Özgün Cicek , Kevin Buchicchio , Thomas Brox

In this paper, we address the problem of forecasting the trajectory of an egocentric camera wearer (ego-person) in crowded spaces. The trajectory forecasting ability learned from the data of different camera wearers walking around in the…

Computer Vision and Pattern Recognition · Computer Science 2022-07-08 Jianing Qiu , Lipeng Chen , Xiao Gu , Frank P. -W. Lo , Ya-Yen Tsai , Jiankai Sun , Jiaqi Liu , Benny Lo

Egocentric interactive world models are essential for augmented reality and embodied AI, where visual generation must respond to user input with low latency, geometric consistency, and long-term stability. We study egocentric interaction…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Yuxi Wang , Wenqi Ouyang , Tianyi Wei , Yi Dong , Zhiqi Shen , Xingang Pan

As robots are increasingly deployed in diverse application domains, enabling robust mobility across different embodiments has become a critical challenge. Classical mobility stacks, though effective on specific platforms, require extensive…

Robotics · Computer Science 2025-10-29 Wei Liu , Huihua Zhao , Chenran Li , Yuchen Deng , Joydeep Biswas , Soha Pouya , Yan Chang

Humans naturally perceive surrounding scenes by unifying sound and sight in a first-person view. Likewise, machines are advanced to approach human intelligence by learning with multisensory inputs from an egocentric perspective. In this…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Chao Huang , Yapeng Tian , Anurag Kumar , Chenliang Xu