English
Related papers

Related papers: Descrip3D: Enhancing Large Language Model-based 3D…

200 papers

Understanding scene contexts is crucial for machines to perform tasks and adapt prior knowledge in unseen or noisy 3D environments. As data-driven learning is intractable to comprehensively encapsulate diverse ranges of layouts and open…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Junho Kim , Gwangtak Bae , Eun Sun Lee , Young Min Kim

Robots are finding wider adoption in human environments, increasing the need for natural human-robot interaction. However, understanding a natural language command requires the robot to infer the intended task and how to decompose it into…

Robotics · Computer Science 2026-02-05 Julia Kuhn , Francesco Verdoja , Tsvetomila Mihaylova , Ville Kyrki

Vision--language models reliably name objects in a scene, but do they represent the 3D layout those objects inhabit? We introduce a 3,034-sample human-curated benchmark targeting three components of spatial understanding: depth-ordered…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Animesh Maheshwari , Divyansh Sahu , Nishit Verma

The vision-based semantic scene completion task aims to predict dense geometric and semantic 3D scene representations from 2D images. However, the presence of dynamic objects in the scene seriously affects the accuracy of the model…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Meng Wang , Fan Wu , Yunchuan Qin , Ruihui Li , Zhuo Tang , Kenli Li

Recent advances in 3D datasets and multimodal models have greatly improved natural language 3D scene understanding. However, most 3D referring segmentation methods do not explicitly represent the observer viewpoint, making spatial relations…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Ayaka Nanri , Klara Reichard , Mert Kiray , Federico Tombari , Benjamin Busam , Asako Kanezaki

Robotic tasks such as planning and navigation require a hierarchical semantic understanding of a scene, which could include multiple floors and rooms. Current methods primarily focus on object segmentation for 3D scene understanding.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Yash Mehan , Kumaraditya Gupta , Rohit Jayanti , Anirudh Govil , Sourav Garg , Madhava Krishna

Identifying objects in an image and their mutual relationships as a scene graph leads to a deep understanding of image content. Despite the recent advancement in deep learning, the detection and labeling of visual object relationships…

Computer Vision and Pattern Recognition · Computer Science 2021-07-13 Rajat Koner , Poulami Sinhamahapatra , Volker Tresp

Automated textual description of remote sensing images is crucial for unlocking their full potential in diverse applications, from environmental monitoring to urban planning and disaster management. However, existing studies in remote…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Kaiyu Li , Zixuan Jiang , Xiangyong Cao , Jiayu Wang , Yuchen Xiao , Deyu Meng , Zhi Wang

Dense visual perception tasks have been constrained by their reliance on predefined categories, limiting their applicability in real-world scenarios where visual concepts are unbounded. While Vision-Language Models (VLMs) like CLIP have…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Junjie Wang , Keyu Chen , Yulin Li , Bin Chen , Hengshuang Zhao , Xiaojuan Qi , Zhuotao Tian

Designing 3D scenes is currently a creative task that requires significant expertise and effort in using complex 3D design interfaces. This effortful design process starts in stark contrast to the easiness with which people can use language…

Graphics · Computer Science 2017-03-02 Angel X. Chang , Mihail Eric , Manolis Savva , Christopher D. Manning

Foundation models have achieved remarkable results in 2D and language tasks like image segmentation, object detection, and visual-language understanding. However, their potential to enrich 3D scene representation learning is largely…

Computer Vision and Pattern Recognition · Computer Science 2023-11-03 Zhimin Chen , Longlong Jing , Yingwei Li , Bing Li

We introduce PAT3D, the first physics-augmented text-to-3D scene generation framework that integrates vision-language models with physics-based simulation to produce physically plausible, simulation-ready, and intersection-free 3D scenes.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Guying Lin , Kemeng Huang , Michael Liu , Ruihan Gao , Hanke Chen , Lyuhao Chen , Beijia Lu , Taku Komura , Yuan Liu , Jun-Yan Zhu , Minchen Li

D scene graphs are an emerging 3D scene representation, that models both the objects present in the scene as well as their relationships. However, learning 3D scene graphs is a challenging task because it requires not only object labels but…

Computer Vision and Pattern Recognition · Computer Science 2023-10-26 Sebastian Koch , Pedro Hermosilla , Narunas Vaskevicius , Mirco Colosi , Timo Ropinski

We address the new problem of language-guided semantic style transfer of 3D indoor scenes. The input is a 3D indoor scene mesh and several phrases that describe the target scene. Firstly, 3D vertex coordinates are mapped to RGB residues by…

Computer Vision and Pattern Recognition · Computer Science 2022-08-17 Bu Jin , Beiwen Tian , Hao Zhao , Guyue Zhou

Grounding object properties and relations in 3D scenes is a prerequisite for a wide range of artificial intelligence tasks, such as visually grounded dialogues and embodied manipulation. However, the variability of the 3D domain induces two…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Joy Hsu , Jiayuan Mao , Jiajun Wu

3D dense captioning stands as a cornerstone in achieving a comprehensive understanding of 3D scenes through natural language. It has recently witnessed remarkable achievements, particularly in indoor settings. However, the exploration of 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-06-07 Bu Jin , Yupeng Zheng , Pengfei Li , Weize Li , Yuhang Zheng , Sujie Hu , Xinyu Liu , Jinwei Zhu , Zhijie Yan , Haiyang Sun , Kun Zhan , Peng Jia , Xiaoxiao Long , Yilun Chen , Hao Zhao

We propose Perceptual Taxonomy, a structured process of scene understanding that first recognizes objects and their spatial configurations, then infers task-relevant properties such as material, affordance, function, and physical attributes…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Jonathan Lee , Xingrui Wang , Jiawei Peng , Luoxin Ye , Zehan Zheng , Tiezheng Zhang , Tao Wang , Wufei Ma , Siyi Chen , Yu-Cheng Chou , Prakhar Kaushik , Alan Yuille

Several modeling domains make use of three-dimensional representations, e.g., the "ball-and-stick" models of molecules. Our generator framework DEViL3D supports the design and implementation of visual 3D languages for such modeling…

Programming Languages · Computer Science 2013-11-21 Jan Wolter

3D scene understanding spans reasoning about free space, object grounding, hypothetical object insertions, complex geometric relationships, and integrating all of these with external tools and data sources. Existing 3D understanding methods…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Sagar Bharadwaj , Ziyong Ma , Anurag Ghosh , Srinivasan Seshan , Anthony Rowe

Open-vocabulary scene understanding is crucial for robotic applications, enabling robots to comprehend complex 3D environmental contexts and supporting various downstream tasks such as navigation and manipulation. However, existing methods…