English
Related papers

Related papers: Chat-3D: Data-efficiently Tuning Large Language Mo…

200 papers

Understanding 3D scenes in open-world settings poses fundamental challenges for vision and robotics, particularly due to the limitations of closed-vocabulary supervision and static annotations. To address this, we propose a unified…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Fei Yu , Quan Deng , Shengeng Tang , Yuehua Li , Lechao Cheng

This paper introduces a novel method for open-vocabulary 3D scene querying in autonomous driving by combining Language Embedded 3D Gaussians with Large Language Models (LLMs). We propose utilizing LLMs to generate both contextually…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Amirhosein Chahe , Lifeng Zhou

Large Vision Language Models (LVLMs) have shown strong capabilities in understanding and analyzing visual scenes across various domains. However, in the context of autonomous driving, their limited comprehension of 3D environments restricts…

Computer Vision and Pattern Recognition · Computer Science 2025-05-02 Jannik Lübberstedt , Esteban Rivera , Nico Uhlemann , Markus Lienkamp

Recent advancements in Large Multimodal Models (LMMs) have greatly enhanced their proficiency in 2D visual understanding tasks, enabling them to effectively process and understand images and videos. However, the development of LMMs with 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Chenming Zhu , Tai Wang , Wenwei Zhang , Jiangmiao Pang , Xihui Liu

Multimodal Large Language Models (MLLMs) have made impressive progress in connecting vision and language, but they still struggle with spatial understanding and viewpoint-aware reasoning. Recent efforts aim to augment the input…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Kevin Qu , Haozhe Qi , Mihai Dusmanu , Mahdi Rad , Rui Wang , Marc Pollefeys

Open-vocabulary scene understanding is crucial for robotic applications, enabling robots to comprehend complex 3D environmental contexts and supporting various downstream tasks such as navigation and manipulation. However, existing methods…

3D spatial understanding is essential in real-world applications such as robotics, autonomous vehicles, virtual reality, and medical imaging. Recently, Large Language Models (LLMs), having demonstrated remarkable success across various…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Jirong Zha , Yuxuan Fan , Xiao Yang , Chen Gao , Xinlei Chen

Traditionally, 3D scene synthesis requires expert knowledge and significant manual effort. Automating this process could greatly benefit fields such as architectural design, robotics simulation, virtual reality, and gaming. Recent…

Computer Vision and Pattern Recognition · Computer Science 2025-06-26 Rui Huang , Guangyao Zhai , Zuria Bauer , Marc Pollefeys , Federico Tombari , Leonidas Guibas , Gao Huang , Francis Engelmann

While multi-modality large language models excel in object-centric or indoor scenarios, scaling them to 3D city-scale environments remains a formidable challenge. To bridge this gap, we propose 3DCity-LLM, a unified framework designed for…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Yiping Chen , Jinpeng Li , Wenyu Ke , Yang Luo , Jie Ouyang , Zhongjie He , Li Liu , Hongchao Fan , Hao Wu

As more applications of large language models (LLMs) for 3D content for immersive environments emerge, it is crucial to study user behaviour to identify interaction patterns and potential barriers to guide the future design of immersive…

Human-Computer Interaction · Computer Science 2026-04-09 Junlong Chen , Jens Grubert , Per Ola Kristensson

3D scene understanding spans reasoning about free space, object grounding, hypothetical object insertions, complex geometric relationships, and integrating all of these with external tools and data sources. Existing 3D understanding methods…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Sagar Bharadwaj , Ziyong Ma , Anurag Ghosh , Srinivasan Seshan , Anthony Rowe

3D Multi-modal Large Language Models (MLLMs) still lag behind their 2D peers, largely because large-scale, high-quality 3D scene-dialogue datasets remain scarce. Prior efforts hinge on expensive human annotation and leave two key…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Siyuan Wei , Chunjie Wang , Xiao Liu , Xiaosheng Yan , Zhishan Zhou , Rui Huang

Large vision-language models (VLMs) have made significant strides in 2D visual understanding tasks, sparking interest in extending these capabilities to 3D scene understanding. However, current 3D VLMs often struggle with robust reasoning…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Ting Huang , Zeyu Zhang , Hao Tang

Previous research has investigated the application of Multimodal Large Language Models (MLLMs) in understanding 3D scenes by interpreting them as videos. These approaches generally depend on comprehensive 3D data inputs, such as point…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Duo Zheng , Shijia Huang , Yanyang Li , Liwei Wang

Understanding scene contexts is crucial for machines to perform tasks and adapt prior knowledge in unseen or noisy 3D environments. As data-driven learning is intractable to comprehensively encapsulate diverse ranges of layouts and open…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Junho Kim , Gwangtak Bae , Eun Sun Lee , Young Min Kim

Enabling Large Language Models (LLMs) to interact with 3D environments is challenging. Existing approaches extract point clouds either from ground truth (GT) geometry or 3D scenes reconstructed by auxiliary models. Text-image aligned 2D…

Computer Vision and Pattern Recognition · Computer Science 2024-04-22 Tao Chu , Pan Zhang , Xiaoyi Dong , Yuhang Zang , Qiong Liu , Jiaqi Wang

The advent of generalist Large Language Models (LLMs) and Large Vision Models (VLMs) have streamlined the construction of semantically enriched maps that can enable robots to ground high-level reasoning and planning into their…

Robotics · Computer Science 2024-11-06 Emilio Olivastri , Jonathan Francis , Alberto Pretto , Niko Sünderhauf , Krishan Rana

The ability to understand and reason the 3D real world is a crucial milestone towards artificial general intelligence. The current common practice is to finetune Large Language Models (LLMs) with 3D data and texts to enable 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Sha Zhang , Di Huang , Jiajun Deng , Shixiang Tang , Wanli Ouyang , Tong He , Yanyong Zhang

Designing 3D indoor layouts is a crucial task with significant applications in virtual reality, interior design, and automated space planning. Existing methods for 3D layout design either rely on diffusion models, which utilize spatial…

Computer Vision and Pattern Recognition · Computer Science 2024-06-07 Yixuan Yang , Junru Lu , Zixiang Zhao , Zhen Luo , James J. Q. Yu , Victor Sanchez , Feng Zheng

Large vision-language models (LVLMs) are increasingly deployed in interactive applications such as virtual and augmented reality, where a first-person (egocentric) view captured by head-mounted cameras serves as key input. While this view…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Insu Lee , Wooje Park , Jaeyun Jang , Minyoung Noh , Kyuhong Shim , Byonghyo Shim