English
Related papers

Related papers: Habitat-Matterport 3D Semantics Dataset

200 papers

To ensure the efficiency of robot autonomy under diverse real-world conditions, a high-quality heterogeneous dataset is essential to benchmark the operating algorithms' performance and robustness. Current benchmarks predominantly focus on…

With the increasing global popularity of self-driving cars, there is an immediate need for challenging real-world datasets for benchmarking and training various computer vision tasks such as 3D object detection. Existing datasets either…

Computer Vision and Pattern Recognition · Computer Science 2019-09-18 Quang-Hieu Pham , Pierre Sevestre , Ramanpreet Singh Pahwa , Huijing Zhan , Chun Ho Pang , Yuda Chen , Armin Mustafa , Vijay Chandrasekhar , Jie Lin

The integration of language and 3D perception is crucial for embodied agents and robots that comprehend and interact with the physical world. While large language models (LLMs) have demonstrated impressive language understanding and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Jianing Yang , Xuweiyi Chen , Nikhil Madaan , Madhavan Iyengar , Shengyi Qian , David F. Fouhey , Joyce Chai

We present ScanNet++, a large-scale dataset that couples together capture of high-quality and commodity-level geometry and color of indoor scenes. Each scene is captured with a high-end laser scanner at sub-millimeter resolution, along with…

Computer Vision and Pattern Recognition · Computer Science 2023-08-23 Chandan Yeshwanth , Yueh-Cheng Liu , Matthias Nießner , Angela Dai

The volumetric representation of human interactions is one of the fundamental domains in the development of immersive media productions and telecommunication applications. Particularly in the context of the rapid advancement of Extended…

Computer Vision and Pattern Recognition · Computer Science 2024-02-15 Fatemeh Ghorbani Lohesara , Davi Rabbouni Freitas , Christine Guillemot , Karen Eguiazarian , Sebastian Knorr

3D scene reconstruction from 2D images is one of the most important tasks in computer graphics. Unfortunately, existing datasets and benchmarks concentrate on idealized synthetic or meticulously captured realistic data. Such benchmarks fail…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Weronika Smolak-Dyżewska , Dawid Malarz , Grzegorz Wilczyński , Rafał Tobiasz , Joanna Waczyńska , Piotr Borycki , Przemysław Spurek

Affordance learning is a complex challenge in many applications, where existing approaches primarily focus on the geometric structures, visual knowledge, and affordance labels of objects to determine interactable regions. However, extending…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Nghia Vu , Tuong Do , Khang Nguyen , Baoru Huang , Nhat Le , Binh Xuan Nguyen , Erman Tjiputra , Quang D. Tran , Ravi Prakash , Te-Chuan Chiu , Anh Nguyen

In this work, we present 3DCoMPaT$^{++}$, a multimodal 2D/3D dataset with 160 million rendered views of more than 10 million stylized 3D shapes carefully annotated at the part-instance level, alongside matching RGB point clouds, 3D textured…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Habib Slim , Xiang Li , Yuchen Li , Mahmoud Ahmed , Mohamed Ayman , Ujjwal Upadhyay , Ahmed Abdelreheem , Arpit Prajapati , Suhail Pothigara , Peter Wonka , Mohamed Elhoseiny

We present HOI4D, a large-scale 4D egocentric dataset with rich annotations, to catalyze the research of category-level human-object interaction. HOI4D consists of 2.4M RGB-D egocentric video frames over 4000 sequences collected by 4…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Yunze Liu , Yun Liu , Che Jiang , Kangbo Lyu , Weikang Wan , Hao Shen , Boqiang Liang , Zhoujie Fu , He Wang , Li Yi

This paper presents a reinforcement learning method for object goal navigation (ObjNav) where an agent navigates in 3D indoor environments to reach a target object based on long-term observations of objects and scenes. To this end, we…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Rui Fukushima , Kei Ota , Asako Kanezaki , Yoko Sasaki , Yusuke Yoshiyasu

Accurate hand pose estimation at joint level has several uses on human-robot interaction, user interfacing and virtual reality applications. Yet, it currently is not a solved problem. The novel deep learning techniques could make a great…

Human-Computer Interaction · Computer Science 2017-07-20 Francisco Gomez-Donoso , Sergio Orts-Escolano , Miguel Cazorla

Omnidirectional images are one of the main sources of information for learning based scene understanding algorithms. However, annotated datasets of omnidirectional images cannot keep the pace of these learning based algorithms development.…

Databases · Computer Science 2024-01-31 Bruno Berenguel-Baeta , Jesus Bermudez-Cameo , Jose J. Guerrero

A bathtub in a library, a sink in an office, a bed in a laundry room -- the counter-intuition suggests that scene provides important prior knowledge for 3D object detection, which instructs to eliminate the ambiguous detection of similar…

Computer Vision and Pattern Recognition · Computer Science 2022-04-13 Yu Zheng , Yueqi Duan , Jiwen Lu , Jie Zhou , Qi Tian

Pixel-level 2D object semantic understanding is an important topic in computer vision and could help machine deeply understand objects (e.g. functionality and affordance) in our daily life. However, most previous methods directly train on…

Computer Vision and Pattern Recognition · Computer Science 2021-11-23 Yang You , Chengkun Li , Yujing Lou , Zhoujun Cheng , Liangwei Li , Lizhuang Ma , Weiming Wang , Cewu Lu

While feed-forward 3D reconstruction models have advanced rapidly, they still exhibit degraded performance on panoramas due to spherical distortions. Moreover, existing panoramic 3D datasets are predominantly collected with 360 cameras…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Jing Ou , Zidong Cao , Yinrui Ren , Zhuoxiao Li , Jinjing Zhu , Tongyan Hua , Shuai Zhang , Hui Xiong , Wufan Zhao

3D semantic occupancy prediction aims to reconstruct the 3D geometry and semantics of the surrounding environment. With dense voxel labels, prior works typically formulate it as a dense segmentation task, independently classifying each…

Graphics · Computer Science 2025-06-06 Wuyang Li , Zhu Yu , Alexandre Alahi

Humans naturally interact with both others and the surrounding multiple objects, engaging in various social activities. However, recent advances in modeling human-object interactions mostly focus on perceiving isolated individuals and…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Juze Zhang , Jingyan Zhang , Zining Song , Zhanhe Shi , Chengfeng Zhao , Ye Shi , Jingyi Yu , Lan Xu , Jingya Wang

Automated 3D CT diagnosis empowers clinicians to make timely, evidence-based decisions by enhancing diagnostic accuracy and workflow efficiency. While multimodal large language models (MLLMs) exhibit promising performance in visual-language…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Yanzhao Shi , Xiaodan Zhang , Junzhong Ji , Haoning Jiang , Chengxin Zheng , Yinong Wang , Liangqiong Qu

Indoor rooms are among the most common use cases in 3D scene understanding. Current state-of-the-art methods for this task are driven by large annotated datasets. Room layouts are especially important, consisting of structural elements in…

Computer Vision and Pattern Recognition · Computer Science 2023-12-22 Denys Rozumnyi , Stefan Popov , Kevis-Kokitsi Maninis , Matthias Nießner , Vittorio Ferrari

We present a benchmark for 3D human whole-body pose estimation, which involves identifying accurate 3D keypoints on the entire human body, including face, hands, body, and feet. Currently, the lack of a fully annotated and accurate 3D…

Computer Vision and Pattern Recognition · Computer Science 2023-09-07 Yue Zhu , Nermin Samet , David Picard
‹ Prev 1 4 5 6 7 8 10 Next ›