English
Related papers

Related papers: ScanQA: 3D Question Answering for Spatial Scene Un…

200 papers

Small object-centric spatial understanding in indoor videos remains a significant challenge for multimodal large language models (MLLMs), despite its practical value for object search and assistive applications. Although existing benchmarks…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Zhiyu Zhou , Peilin Liu , Ruoxuan Zhang , Luyang Zhang , Cheng Zhang , Hongxia Xie , Wen-Huang Cheng

Open-vocabulary 3D scene understanding presents a significant challenge in computer vision, with wide-ranging applications in embodied agents and augmented reality systems. Existing methods adopt neurel rendering methods as 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Jun Guo , Xiaojian Ma , Yue Fan , Huaping Liu , Qing Li

Recent advancements in multi-modal large language models (MLLMs) have shown strong potential for 3D scene understanding. However, existing methods struggle with fine-grained object grounding and contextual reasoning, limiting their ability…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Haifeng Huang , Yilun Chen , Zehan Wang , Jiangmiao Pang , Zhou Zhao

Spatial reasoning in large-scale 3D environments such as warehouses remains a significant challenge for vision-language systems due to scene clutter, occlusions, and the need for precise spatial understanding. Existing models often struggle…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Tanner Muturi , Blessing Agyei Kyem , Joshua Kofi Asamoah , Neema Jakisa Owor , Richard Dyzinela , Andrews Danyo , Yaw Adu-Gyamfi , Armstrong Aboah

3D Large Language Models (LLMs) leveraging spatial information in point clouds for 3D spatial reasoning attract great attention. Despite some promising results, the advantages of point clouds over other modalities remain unclear. Moreover,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Weichen Zhang , Ruiying Peng , Xin Zeng , Jianjie Fang , Ziyou Wang , Kaiyuan Li , Heng Dong , Wei Li , Chen Gao , Xin Wang , Xinlei Chen , Yong Li

Prior studies on 3D scene understanding have primarily developed specialized models for specific tasks or required task-specific fine-tuning. In this study, we propose Grounded 3D-LLM, which explores the potential of 3D large multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Yilun Chen , Shuai Yang , Haifeng Huang , Tai Wang , Runsen Xu , Ruiyuan Lyu , Dahua Lin , Jiangmiao Pang

3D understanding is a key capability for real-world AI assistance. High-quality data plays an important role in driving the development of the 3D understanding community. Current 3D scene understanding datasets often provide geometric and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Zirui Wang , Tao Zhang

Document Visual Question Answering (DocVQA) requires models to jointly understand textual semantics, spatial layout, and visual features. Current methods struggle with explicit spatial relationship modeling, inefficiency with…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Ahmad Mohammadshirazi , Pinaki Prasad Guha Neogi , Dheeraj Kulshrestha , Rajiv Ramnath

Complex 3D scene understanding has gained increasing attention, with scene encoding strategies playing a crucial role in this success. However, the optimal scene encoding strategies for various scenarios remain unclear, particularly…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Yunze Man , Shuhong Zheng , Zhipeng Bao , Martial Hebert , Liang-Yan Gui , Yu-Xiong Wang

One major goal of vision is to infer physical models of objects, surfaces, and their layout from sensors. In this paper, we aim to interpret indoor scenes from one RGBD image. Our representation encodes the layout of walls, which must…

Computer Vision and Pattern Recognition · Computer Science 2017-08-21 Ruiqi Guo , Chuhang Zou , Derek Hoiem

Query-based 3D object detection methods using multi-view images often struggle to efficiently leverage dynamic multi-scale information, e.g., the relationship between the object features and the geometric of the queries are not sufficiently…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Mingxi Pang , Dingheng Wang , Zekun Li , Zhenping Sun , Bo Wang , Zhihang Wang , Zhao-Xu Yang

The rapid advancement of Multimodal Large Language Models (MLLMs) has significantly impacted various multimodal tasks. However, these models face challenges in tasks that require spatial understanding within 3D environments. Efforts to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Duo Zheng , Shijia Huang , Liwei Wang

3D object understanding and generation methods produce impressive results, yet they often overlook a pervasive source of information in real-world scenes: repeated objects. We introduce the task of lookalike object detection in indoor…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Chandan Yeshwanth , Angela Dai

The ideal form of Visual Question Answering requires understanding, grounding and reasoning in the joint space of vision and language and serves as a proxy for the AI task of scene understanding. However, most existing VQA benchmarks are…

Computer Vision and Pattern Recognition · Computer Science 2023-03-07 Kang Chen , Xiangqian Wu

Recent studies on dense captioning and visual grounding in 3D have achieved impressive results. Despite developments in both areas, the limited amount of available 3D vision-language data causes overfitting issues for 3D visual grounding…

Computer Vision and Pattern Recognition · Computer Science 2022-07-25 Dave Zhenyu Chen , Qirui Wu , Matthias Nießner , Angel X. Chang

We present a novel approach to reconstructing lightweight, CAD-based representations of scanned 3D environments from commodity RGB-D sensors. Our key idea is to jointly optimize for both CAD model alignments as well as layout estimations of…

Computer Vision and Pattern Recognition · Computer Science 2020-03-31 Armen Avetisyan , Tatiana Khanova , Christopher Choy , Denver Dash , Angela Dai , Matthias Nießner

Embodied Question Answering (EQA) connects perception, reasoning, and interaction within embodied environments. However, existing datasets and benchmarks remain fragmented, each focusing on a limited subset of reasoning skills such as…

Robotics · Computer Science 2026-05-26 Xicheng Gong , Qiwei Li , Peiran Xu , Yadong Mu

Creating machines capable of understanding the world in 3D is essential in assisting designers that build and edit 3D environments and robots navigating and interacting within a three-dimensional space. Inspired by advances in language and…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Aadarsh Sahoo , Vansh Tibrewal , Georgia Gkioxari

Visual question answering (VQA) has recently been introduced to remote sensing to make information extraction from overhead imagery more accessible to everyone. VQA considers a question (in natural language, therefore easy to formulate)…

Computer Vision and Pattern Recognition · Computer Science 2021-09-27 Christel Chappuis , Sylvain Lobry , Benjamin Kellenberger , Bertrand Le Saux , Devis Tuia

This paper presents stacked attention networks (SANs) that learn to answer natural language questions from images. SANs use semantic representation of a question as query to search for the regions in an image that are related to the answer.…

Machine Learning · Computer Science 2016-01-27 Zichao Yang , Xiaodong He , Jianfeng Gao , Li Deng , Alex Smola