English
Related papers

Related papers: PerLA: Perceptive 3D Language Assistant

200 papers

3D Large Language Models (LLMs) leveraging spatial information in point clouds for 3D spatial reasoning attract great attention. Despite some promising results, the advantages of point clouds over other modalities remain unclear. Moreover,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Weichen Zhang , Ruiying Peng , Xin Zeng , Jianjie Fang , Ziyou Wang , Kaiyuan Li , Heng Dong , Wei Li , Chen Gao , Xin Wang , Xinlei Chen , Yong Li

Affordance understanding, the task of identifying actionable regions on 3D objects, plays a vital role in allowing robotic systems to engage with and operate within the physical world. Although Visual Language Models (VLMs) have excelled in…

3D dense captioning, as an emerging vision-language task, aims to identify and locate each object from a set of point clouds and generate a distinctive natural language sentence for describing each located object. However, the existing…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Yufeng Zhong , Long Xu , Jiebo Luo , Lin Ma

Vision-language models (VLMs), such as CLIP and ALIGN, are generally trained on datasets consisting of image-caption pairs obtained from the web. However, real-world multimodal datasets, such as healthcare data, are significantly more…

Computer Vision and Pattern Recognition · Computer Science 2023-08-23 Maya Varma , Jean-Benoit Delbrouck , Sarah Hooper , Akshay Chaudhari , Curtis Langlotz

We introduce a relevant yet challenging problem named Personalized Dictionary Learning (PerDL), where the goal is to learn sparse linear representations from heterogeneous datasets that share some commonality. In PerDL, we model each…

Machine Learning · Computer Science 2023-05-25 Geyu Liang , Naichen Shi , Raed Al Kontar , Salar Fattahi

Large 2D vision-language models (2D-LLMs) have gained significant attention by bridging Large Language Models (LLMs) with images using a simple projector. Inspired by their success, large 3D point cloud-language models (3D-LLMs) also…

Computer Vision and Pattern Recognition · Computer Science 2024-05-03 Yuan Tang , Xu Han , Xianzhi Li , Qiao Yu , Yixue Hao , Long Hu , Min Chen

Place recognition gives a SLAM system the ability to correct cumulative errors. Unlike images that contain rich texture features, point clouds are almost pure geometric information which makes place recognition based on point clouds…

Computer Vision and Pattern Recognition · Computer Science 2021-07-13 Lin Li , Xin Kong , Xiangrui Zhao , Tianxin Huang , Yong Liu

Recent advancements in multimodal large language models (LLMs) have demonstrated significant potential across various domains, particularly in concept reasoning. However, their applications in understanding 3D environments remain limited,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Kuan-Chih Huang , Xiangtai Li , Lu Qi , Shuicheng Yan , Ming-Hsuan Yang

Self-supervised learning (SSL) on 3D point clouds has the potential to learn feature representations that can transfer to diverse sensors and multiple downstream perception tasks. However, recent SSL approaches fail to define pretext tasks…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Barza Nisar , Steven L. Waslander

Personality detection automatically identifies an individual's personality from various data sources, such as social media texts. However, as the parameter scale of language models continues to grow, the computational cost becomes…

Computation and Language · Computer Science 2025-04-09 Lingzhi Shen , Yunfei Long , Xiaohao Cai , Guanming Chen , Imran Razzak , Shoaib Jameel

Multi-modal Large Language Models (MLLMs) exhibit impressive capabilities in 2D tasks, yet encounter challenges in discerning the spatial positions, interrelations, and causal logic in scenes when transitioning from 2D to 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-01-15 Haomiao Xiong , Yunzhi Zhuge , Jiawen Zhu , Lu Zhang , Huchuan Lu

Learning to ground natural language queries to target objects or regions in 3D point clouds is quite essential for 3D scene understanding. Nevertheless, existing 3D visual grounding approaches require a substantial number of bounding box…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Xiaoxu Xu , Yitian Yuan , Qiudan Zhang , Wenhui Wu , Zequn Jie , Lin Ma , Xu Wang

Enabling Large Language Models (LLMs) to interact with 3D environments is challenging. Existing approaches extract point clouds either from ground truth (GT) geometry or 3D scenes reconstructed by auxiliary models. Text-image aligned 2D…

Computer Vision and Pattern Recognition · Computer Science 2024-04-22 Tao Chu , Pan Zhang , Xiaoyi Dong , Yuhang Zang , Qiong Liu , Jiaqi Wang

While multi-modality large language models excel in object-centric or indoor scenarios, scaling them to 3D city-scale environments remains a formidable challenge. To bridge this gap, we propose 3DCity-LLM, a unified framework designed for…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Yiping Chen , Jinpeng Li , Wenyu Ke , Yang Luo , Jie Ouyang , Zhongjie He , Li Liu , Hongchao Fan , Hao Wu

Existing state-of-the-art 3D point cloud understanding methods merely perform well in a fully supervised manner. To the best of our knowledge, there exists no unified framework that simultaneously solves the downstream high-level…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Kangcheng Liu

Multimodal Large Language Models (MLLMs) are widely used for visual perception, understanding, and reasoning. However, long video processing and precise moment retrieval remain challenging due to LLMs' limited context size and coarse frame…

Computer Vision and Pattern Recognition · Computer Science 2024-11-25 Weiheng Lu , Jian Li , An Yu , Ming-Ching Chang , Shengpeng Ji , Min Xia

We tackle the problem of localizing 3D point cloud submaps using complex and diverse natural language descriptions, and present Text2Loc++, a novel neural network designed for effective cross-modal alignment between language and point…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Yan Xia , Letian Shi , Yilin Di , Joao F. Henriques , Daniel Cremers

As three-dimensional acquisition technologies like LiDAR cameras advance, the need for efficient transmission of 3D point clouds is becoming increasingly important. In this paper, we present a novel semantic communication (SemCom) approach…

Emerging Technologies · Computer Science 2025-05-13 Shangzhuo Xie , Qianqian Yang , Yuyi Sun , Tianxiao Han , Zhaohui Yang , Zhiguo Shi

Point clouds are a key modality used for perception in autonomous vehicles, providing the means for a robust geometric understanding of the surrounding environment. However despite the sensor outputs from autonomous vehicles being naturally…

Computer Vision and Pattern Recognition · Computer Science 2021-12-03 Joshua Knights , Peyman Moghadam , Clinton Fookes , Sridha Sridharan

3D Multimodal Large Language Models (MLLMs) have recently made substantial advancements. However, their potential remains untapped, primarily due to the limited quantity and suboptimal quality of 3D datasets. Current approaches attempt to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Zilu Guo , Hongbin Lin , Zhihao Yuan , Chaoda Zheng , Pengshuo Qiu , Dongzhi Jiang , Renrui Zhang , Chun-Mei Feng , Zhen Li