English
Related papers

Related papers: Habitat-Matterport 3D Semantics Dataset

200 papers

A deep understanding of the physical world is a central goal for embodied AI and realistic simulation. While current models excel at capturing an object's surface geometry and appearance, they largely neglect its internal physical…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Jingxuan Zhang , Tianqi Yu , Yatu Zhang , Jinze Wu , Kaixin Yao , Jingyang Liu , Yuyao Zhang , Jiayuan Gu , Jingyi Yu

Semantic segmentation has emerged as a pivotal area of study in computer vision, offering profound implications for scene understanding and elevating human-machine interactions across various domains. While 2D semantic segmentation has…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Aditya Krishnan , Jayneel Vora , Prasant Mohapatra

Dense 3D semantic occupancy perception is critical for mobile robots operating in pedestrian-rich environments, yet it remains underexplored compared to its application in autonomous driving. To address this gap, we present MobileOcc, a…

Robotics · Computer Science 2025-11-24 Junseo Kim , Guido Dumont , Xinyu Gao , Gang Chen , Holger Caesar , Javier Alonso-Mora

Prior studies on 3D scene understanding have primarily developed specialized models for specific tasks or required task-specific fine-tuning. In this study, we propose Grounded 3D-LLM, which explores the potential of 3D large multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Yilun Chen , Shuai Yang , Haifeng Huang , Tai Wang , Runsen Xu , Ruiyuan Lyu , Dahua Lin , Jiangmiao Pang

The use of RGB-D information for salient object detection has been extensively explored in recent years. However, relatively few efforts have been put towards modeling salient object detection in real-world human activity scenes with RGBD.…

Computer Vision and Pattern Recognition · Computer Science 2024-02-21 Deng-Ping Fan , Zheng Lin , Jia-Xing Zhao , Yun Liu , Zhao Zhang , Qibin Hou , Menglong Zhu , Ming-Ming Cheng

Large Multimodal Models (LMMs) have become a pivotal research focus in deep learning, demonstrating remarkable capabilities in 3D scene understanding. However, current 3D LMMs employing thousands of spatial tokens for multimodal reasoning…

Graphics · Computer Science 2025-05-20 Kai Zhang , Xingyu Chen , Xiaofeng Zhang

In this paper, we present the first large-scale dataset for semantic Segmentation of Underwater IMagery (SUIM). It contains over 1500 images with pixel annotations for eight object categories: fish (vertebrates), reefs (invertebrates),…

Computer Vision and Pattern Recognition · Computer Science 2020-09-15 Md Jahidul Islam , Chelsey Edge , Yuyang Xiao , Peigen Luo , Muntaqim Mehtaz , Christopher Morse , Sadman Sakib Enan , Junaed Sattar

In this paper, we introduce a novel benchmark designed to propel the advancement of general-purpose, large-scale 3D vision models for remote sensing imagery. While several datasets have been proposed within the realm of remote sensing, many…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Jiayu Wang , Ruizhi Wang , Jie Song , Haofei Zhang , Mingli Song , Zunlei Feng , Li Sun

Robotic manipulation and navigation are fundamental capabilities of embodied intelligence, enabling effective robot interactions with the physical world. Achieving these capabilities requires a cohesive understanding of the environment,…

Robotics · Computer Science 2025-11-18 Xiaoshuai Hao , Yingbo Tang , Lingfeng Zhang , Yanbiao Ma , Yunfeng Diao , Ziyu Jia , Wenbo Ding , Hangjun Ye , Long Chen

Even though a significant amount of work has been done to increase the safety of transportation networks, accidents still occur regularly. They must be understood as unavoidable and sporadic outcomes of traffic networks. No public dataset…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Walter Zimmer , Ross Greer , Daniel Lehmberg , Marc Pavel , Holger Caesar , Xingcheng Zhou , Ahmed Ghita , Mohan Trivedi , Rui Song , Hu Cao , Akshay Gopalkrishnan , Alois C. Knoll

Evaluating the performance of Multi-modal Large Language Models (MLLMs), integrating both point cloud and language, presents significant challenges. The lack of a comprehensive assessment hampers determining whether these models truly…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 Junjie Zhang , Tianci Hu , Xiaoshui Huang , Yongshun Gong , Dan Zeng

Understanding human behaviour in crowded indoor environments is central to surveillance, smart buildings, and human-robot interaction, yet existing datasets rarely capture real-world indoor complexity at scale. We introduce IndoorCrowd, a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Sebastian-Ion Nae , Radu Moldoveanu , Alexandra Stefania Ghita , Adina Magda Florea

360 video captures the complete surrounding scenes with the ultra-large field of view of 360X180. This makes 360 scene understanding tasks, eg, segmentation and tracking, crucial for appications, such as autonomous driving, robotics. With…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Weiming Zhang , Dingwen Xiao , Aobotao Dai , Yexin Liu , Tianbo Pan , Shiqi Wen , Lei Chen , Lin Wang

Recognizing scenes and objects in 3D from a single image is a longstanding goal of computer vision with applications in robotics and AR/VR. For 2D recognition, large datasets and scalable solutions have led to unprecedented advances. In 3D,…

Computer Vision and Pattern Recognition · Computer Science 2023-03-27 Garrick Brazil , Abhinav Kumar , Julian Straub , Nikhila Ravi , Justin Johnson , Georgia Gkioxari

Latent representations learned by neural networks often exhibit semantic structure, where concept similarity is reflected by geometric proximity in embedding space. However, comparing such spaces across models remains difficult: changes in…

In vision-and-language navigation (VLN), an embodied agent is required to navigate in realistic 3D environments following natural language instructions. One major bottleneck for existing VLN approaches is the lack of sufficient training…

Computer Vision and Pattern Recognition · Computer Science 2022-08-26 Shizhe Chen , Pierre-Louis Guhur , Makarand Tapaswi , Cordelia Schmid , Ivan Laptev

Deployment of deep learning models in robotics as sensory information extractors can be a daunting task to handle, even using generic GPU cards. Here, we address three of its most prominent hurdles, namely, i) the adaptation of a single…

Computer Vision and Pattern Recognition · Computer Science 2019-02-28 Vladimir Nekrasov , Thanuja Dharmasiri , Andrew Spek , Tom Drummond , Chunhua Shen , Ian Reid

The existence of variable factors within the environment can cause a decline in camera localization accuracy, as it violates the fundamental assumption of a static environment in Simultaneous Localization and Mapping (SLAM) algorithms.…

Robotics · Computer Science 2023-10-11 Ghanta Sai Krishna , Kundrapu Supriya , Sabur Baidya

Moving around in the world is naturally a multisensory experience, but today's embodied agents are deaf---restricted to solely their visual perception of the environment. We introduce audio-visual navigation for complex, acoustically and…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Changan Chen , Unnat Jain , Carl Schissler , Sebastia Vicenc Amengual Gari , Ziad Al-Halah , Vamsi Krishna Ithapu , Philip Robinson , Kristen Grauman

Accurate 3D reconstruction of objects with reflective, transparent, or low-texture surfaces still remains notoriously challenging. Such materials often violate key assumptions in multi-view reconstruction pipelines, such as photometric…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Zhicheng Liang , Haoyi Yu , Boyan Li , Dayou Zhang , Zijian Cao , Tianyi Gong , Junhua Liu , Shuguang Cui , Fangxin Wang