English
Related papers

Related papers: Chat-3D: Data-efficiently Tuning Large Language Mo…

200 papers

Despite recent advances in multimodal content generation enabled by vision-language models (VLMs), their ability to reason about and generate structured 3D scenes remains largely underexplored. This limitation constrains their utility in…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Xinhang Liu , Yu-Wing Tai , Chi-Keung Tang

Large Language Models (LLMs) represent formidable tools for sequence modeling, boasting an innate capacity for general pattern recognition. Nevertheless, their broader spatial reasoning capabilities, especially applied to numerical…

Robotics · Computer Science 2023-12-05 Manasi Sharma

Open-set perception in complex traffic environments poses a critical challenge for autonomous driving systems, particularly in identifying previously unseen object categories, which is vital for ensuring safety. Visual Language Models…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Fuhao Chang , Shuxin Li , Yabei Li , Lei He

SpatialLM is a large language model designed to process 3D point cloud data and generate structured 3D scene understanding outputs. These outputs include architectural elements like walls, doors, windows, and oriented object boxes with…

Computer Vision and Pattern Recognition · Computer Science 2025-11-06 Yongsen Mao , Junhao Zhong , Chuan Fang , Jia Zheng , Rui Tang , Hao Zhu , Ping Tan , Zihan Zhou

Large language models have emerged as a promising approach towards achieving general-purpose AI agents. The thriving open-source LLM community has greatly accelerated the development of agents that support human-machine dialogue interaction…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Zhenfei Yin , Jiong Wang , Jianjian Cao , Zhelun Shi , Dingning Liu , Mukai Li , Lu Sheng , Lei Bai , Xiaoshui Huang , Zhiyong Wang , Jing Shao , Wanli Ouyang

We propose UniSeg3D, a unified 3D scene understanding framework that achieves panoptic, semantic, instance, interactive, referring, and open-vocabulary segmentation tasks within a single model. Most previous 3D segmentation approaches are…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 Wei Xu , Chunsheng Shi , Sifan Tu , Xin Zhou , Dingkang Liang , Xiang Bai

We introduce FaceGPT, a self-supervised learning framework for Large Vision-Language Models (VLMs) to reason about 3D human faces from images and text. Typical 3D face reconstruction methods are specialized algorithms that lack semantic…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Haoran Wang , Mohit Mendiratta , Christian Theobalt , Adam Kortylewski

Recently, Large Language Models (LLMs) have achieved significant success, prompting increased interest in expanding their generative capabilities beyond general text into domain-specific areas. This study investigates the generation of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Jiahao Li , Weijian Ma , Xueyang Li , Yunzhong Lou , Guichun Zhou , Xiangdong Zhou

Enabling Large Language Models (LLMs) to understand the 3D physical world is an emerging yet challenging research direction. Current strategies for processing point clouds typically downsample the scene or divide it into smaller parts for…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Guofeng Mei , Wei Lin , Luigi Riz , Yujiao Wu , Fabio Poiesi , Yiming Wang

Recently, the powerful text-to-image capabilities of ChatGPT-4o have led to growing appreciation for native multimodal large language models. However, its multimodal capabilities remain confined to images and text. Yet beyond images, the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Junliang Ye , Zhengyi Wang , Ruowen Zhao , Shenghao Xie , Jun Zhu

Multimodal Large Language Models (MLLMs) struggle with accurately capturing camera-object relations, especially for object orientation, camera viewpoint, and camera shots. This stems from the fact that existing MLLMs are trained on images…

Dialogue data has been a key source for understanding learning processes, offering critical insights into how students engage in collaborative discussions and how these interactions shape their knowledge construction. The advent of Large…

Computation and Language · Computer Science 2025-04-29 Ying Na , Shihui Feng

Recently, large language models (LLMs) have been explored widely for 3D scene understanding. Among them, training-free approaches are gaining attention for their flexibility and generalization over training-based methods. However, they…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Haida Feng , Hao Wei , Zewen Xu , Haolin Wang , Chade Li , Yihong Wu

This paper addresses the task of large-scale 3D scene reconstruction from long video sequences. Recent feed-forward reconstruction models have shown promising results by directly regressing 3D geometry from RGB images without explicit 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Tao Xie , Peishan Yang , Yudong Jin , Yingfeng Cai , Wei Yin , Weiqiang Ren , Qian Zhang , Wei Hua , Sida Peng , Xiaoyang Guo , Xiaowei Zhou

Designing high-quality indoor 3D scenes is important in many practical applications, such as room planning or game development. Conventionally, this has been a time-consuming process which requires both artistic skill and familiarity with…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Başak Melis Öcal , Maxim Tatarchenko , Sezer Karaoglu , Theo Gevers

In recent years, vision language pre-training frameworks have made significant progress in natural language processing and computer vision, achieving remarkable performance improvement on various downstream tasks. However, when extended to…

Computer Vision and Pattern Recognition · Computer Science 2023-05-19 Taolin Zhang , Sunan He , Dai Tao , Bin Chen , Zhi Wang , Shu-Tao Xia

Safety-critical 3D scene understanding tasks necessitate not only accurate but also confident predictions from 3D perception models. This study introduces Calib3D, a pioneering effort to benchmark and scrutinize the reliability of 3D scene…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Lingdong Kong , Xiang Xu , Jun Cen , Wenwei Zhang , Liang Pan , Kai Chen , Ziwei Liu

3D scene understanding has been transformed by open-vocabulary language models that enable interaction via natural language. However, at present the evaluation of these representations is limited to datasets with closed-set semantics that…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Christina Kassab , Sacha Morin , Martin Büchner , Matías Mattamala , Kumaraditya Gupta , Abhinav Valada , Liam Paull , Maurice Fallon

Large language models (LLMs) and their variants have shown extraordinary efficacy across numerous downstream natural language processing (NLP) tasks, which has presented a new vision for the development of NLP. Despite their remarkable…

Computation and Language · Computer Science 2024-01-18 Yazhou Zhang , Mengyao Wang , Youxi Wu , Prayag Tiwari , Qiuchi Li , Benyou Wang , Jing Qin

Training a 3D scene understanding model requires complicated human annotations, which are laborious to collect and result in a model only encoding close-set object semantics. In contrast, vision-language pre-training models (e.g., CLIP)…

Computer Vision and Pattern Recognition · Computer Science 2023-03-17 Junbo Zhang , Runpei Dong , Kaisheng Ma