English
Related papers

Related papers: InHabit: Leveraging Image Foundation Models for Sc…

200 papers

Embodied agents operating in human spaces must be able to master how their environment works: what objects can the agent use, and how can it use them? We introduce a reinforcement learning approach for exploration for interaction, whereby…

Computer Vision and Pattern Recognition · Computer Science 2020-10-20 Tushar Nagarajan , Kristen Grauman

3D simulated environments play a critical role in Embodied AI, but their creation requires expertise and extensive manual effort, restricting their diversity and scope. To mitigate this limitation, we present Holodeck, a system that…

Understanding and forecasting the scene evolutions deeply affect the exploration and decision of embodied agents. While traditional methods simulate scene evolutions through trajectory prediction of potential instances, current works use…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Zhang Zhang , Qiang Zhang , Wei Cui , Shuai Shi , Yijie Guo , Gang Han , Wen Zhao , Jingkai Sun , Jiahang Cao , Jiaxu Wang , Hao Cheng , Xiaozhu Ju , Zhengping Che , Renjing Xu , Jian Tang

We introduce Replica, a dataset of 18 highly photo-realistic 3D indoor scene reconstructions at room and building scale. Each scene consists of a dense mesh, high-resolution high-dynamic-range (HDR) textures, per-primitive semantic class…

3D human pose estimation from a single image is a challenging problem, especially for in-the-wild settings due to the lack of 3D annotated data. We propose two anatomically inspired loss functions and use them with a weakly-supervised…

Computer Vision and Pattern Recognition · Computer Science 2018-07-05 Rishabh Dabral , Anurag Mundhada , Uday Kusupati , Safeer Afaque , Abhishek Sharma , Arjun Jain

Robotic grasping of house-hold objects has made remarkable progress in recent years. Yet, human grasps are still difficult to synthesize realistically. There are several key reasons: (1) the human hand has many degrees of freedom (more than…

Computer Vision and Pattern Recognition · Computer Science 2020-11-30 Korrawe Karunratanakul , Jinlong Yang , Yan Zhang , Michael Black , Krikamol Muandet , Siyu Tang

3D Human Motion Indexing and Retrieval is an interesting problem due to the rise of several data-driven applications aimed at analyzing and/or re-utilizing 3D human skeletal data, such as data-driven animation, analysis of sports…

Computer Vision and Pattern Recognition · Computer Science 2019-12-11 Neeraj Battan , Abbhinav Venkat , Avinash Sharma

Real-world human-built environments are highly dynamic, involving multiple humans and their complex interactions with surrounding objects. While 3D geometry modeling of such scenes is crucial for applications like AR/VR, gaming, and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Sandika Biswas , Qianyi Wu , Biplab Banerjee , Hamid Rezatofighi

The generation of 3D clothed humans has attracted increasing attention in recent years. However, existing work cannot generate layered high-quality 3D humans with consistent body structures. As a result, these methods are unable to…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Yi Wang , Jian Ma , Ruizhi Shao , Qiao Feng , Yu-Kun Lai , Yebin Liu , Kun Li

Human matting, high quality extraction of humans from natural images, is crucial for a wide variety of applications. Since the matting problem is severely under-constrained, most previous methods require user interactions to take user…

Computer Vision and Pattern Recognition · Computer Science 2018-09-19 Quan Chen , Tiezheng Ge , Yanyu Xu , Zhiqiang Zhang , Xinxin Yang , Kun Gai

We propose a novel representation of virtual humans for highly realistic real-time animation and rendering in 3D applications. We learn pose dependent appearance and geometry from highly accurate dynamic mesh sequences obtained from…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Wieland Morgenstern , Milena T. Bagdasarian , Anna Hilsmann , Peter Eisert

Despite large-scale pretraining endowing models with language and vision reasoning capabilities, improving their spatial reasoning capability remains challenging due to the lack of data grounded in the 3D world. While it is possible for…

Real-world scenes, such as those in ScanNet, are difficult to capture, with highly limited data available. Generating realistic scenes with varied object poses remains an open and challenging task. In this work, we propose FactoredScenes, a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Joy Hsu , Emily Jin , Jiajun Wu , Niloy J. Mitra

We propose a novel system for robot-to-human object handover that emulates human coworker interactions. Unlike most existing studies that focus primarily on grasping strategies and motion planning, our system focus on 1. inferring human…

Robotics · Computer Science 2025-03-06 Hanxin Zhang , Abdulqader Dhafer , Zhou Daniel Hao , Hongbiao Dong

Future robots are envisioned as versatile systems capable of performing a variety of household tasks. The big question remains, how can we bridge the embodiment gap while minimizing physical robot learning, which fundamentally does not…

Robotics · Computer Science 2025-03-31 Hanzhi Chen , Boyang Sun , Anran Zhang , Marc Pollefeys , Stefan Leutenegger

The creation of complex 3D scenes tailored to user specifications has been a tedious and challenging task with traditional 3D modeling tools. Although some pioneering methods have achieved automatic text-to-3D generation, they are generally…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Xiuyu Yang , Yunze Man , Jun-Kun Chen , Yu-Xiong Wang

In this era, the success of large language models and text-to-image models can be attributed to the driving force of large-scale datasets. However, in the realm of 3D vision, while remarkable progress has been made with models trained on…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Zhangyang Xiong , Chenghong Li , Kenkun Liu , Hongjie Liao , Jianqiao Hu , Junyi Zhu , Shuliang Ning , Lingteng Qiu , Chongjie Wang , Shijie Wang , Shuguang Cui , Xiaoguang Han

The ability to forecast human-environment collisions from egocentric observations is vital to enable collision avoidance in applications such as VR, AR, and wearable assistive robotics. In this work, we introduce the challenging problem of…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Boxiao Pan , Bokui Shen , Davis Rempe , Despoina Paschalidou , Kaichun Mo , Yanchao Yang , Leonidas J. Guibas

Hatching is a common method used by artists to accentuate the third dimension of a sketch, and to illuminate the scene. Our system SHAD3S attempts to compete with a human at hatching generic three-dimensional (3D) shapes, and also tries to…

Graphics · Computer Science 2021-09-07 Raghav B. Venkataramaiyer , Abhishek Joshi , Saisha Narang , Vinay P. Namboodiri

We present the first approach to build hierarchical task-driven 3D scene graphs of arbitrary indoor or outdoor environments using an uncalibrated monocular camera in real-time. We leverage geometric foundation models to estimate geometric…

Robotics · Computer Science 2026-05-26 Dominic Maggio , Nicolas Gorlo , Luca Carlone