中文
相关论文

相关论文: CL3DOR: Contrastive Learning for 3D Large Multimod…

200 篇论文

3D object classification has attracted appealing attentions in academic researches and industrial applications. However, most existing methods need to access the training data of past 3D object classes when facing the common real-world…

计算机视觉与模式识别 · 计算机科学 2020-12-17 Jiahua Dong , Yang Cong , Gan Sun , Bingtao Ma , Lichen Wang

Despite encouraging progress in 3D scene understanding, it remains challenging to develop an effective Large Multi-modal Model (LMM) that is capable of understanding and reasoning in complex 3D environments. Most previous methods typically…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Hanxun Yu , Wentong Li , Song Wang , Junbo Chen , Jianke Zhu

Point cloud data plays an essential role in robotics and self-driving applications. Yet, annotating point cloud data is time-consuming and nontrivial while they enable learning discriminative 3D representations that empower downstream…

计算机视觉与模式识别 · 计算机科学 2023-03-01 Srikanth Malla , Yi-Ting Chen

In recent years, vision language pre-training frameworks have made significant progress in natural language processing and computer vision, achieving remarkable performance improvement on various downstream tasks. However, when extended to…

计算机视觉与模式识别 · 计算机科学 2023-05-19 Taolin Zhang , Sunan He , Dai Tao , Bin Chen , Zhi Wang , Shu-Tao Xia

Cross-modal 3D retrieval is a critical yet challenging task, aiming to achieve bi-directional retrieval between 3D and text modalities. Current methods predominantly rely on a certain 3D representation (e.g., point cloud), with few…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Junlong Ren , Hao Wang

Current point-cloud detection methods have difficulty detecting the open-vocabulary objects in the real world, due to their limited generalization capability. Moreover, it is extremely laborious and expensive to collect and fully annotate a…

计算机视觉与模式识别 · 计算机科学 2022-07-06 Yuheng Lu , Chenfeng Xu , Xiaobao Wei , Xiaodong Xie , Masayoshi Tomizuka , Kurt Keutzer , Shanghang Zhang

Recent advancements in multimodal large language models (LLMs) have demonstrated significant potential across various domains, particularly in concept reasoning. However, their applications in understanding 3D environments remain limited,…

计算机视觉与模式识别 · 计算机科学 2025-02-11 Kuan-Chih Huang , Xiangtai Li , Lu Qi , Shuicheng Yan , Ming-Hsuan Yang

Multimodal models, such as the Contrastive Language-Image Pre-training (CLIP) model, have demonstrated remarkable success in aligning visual and linguistic representations. However, these models exhibit limitations when applied to…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Hiroshi Sasaki

The remarkable potential of multi-modal large language models (MLLMs) in comprehending both vision and language information has been widely acknowledged. However, the scarcity of 3D scenes-language pairs in comparison to their 2D…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Zeju Li , Chao Zhang , Xiaoyan Wang , Ruilong Ren , Yifan Xu , Ruifei Ma , Xiangde Liu

Multi-modal large language models (MLLMs) have shown incredible capabilities in a variety of 2D vision and language tasks. We extend MLLMs' perceptual capabilities to ground and reason about images in 3-dimensional space. To that end, we…

计算机视觉与模式识别 · 计算机科学 2024-05-07 Jang Hyun Cho , Boris Ivanovic , Yulong Cao , Edward Schmerling , Yue Wang , Xinshuo Weng , Boyi Li , Yurong You , Philipp Krähenbühl , Yan Wang , Marco Pavone

We introduce C3LLM (Conditioned-on-Three-Modalities Large Language Models), a novel framework combining three tasks of video-to-audio, audio-to-text, and text-to-audio together. C3LLM adapts the Large Language Model (LLM) structure as a…

人工智能 · 计算机科学 2024-05-28 Zixuan Wang , Qinkai Duan , Yu-Wing Tai , Chi-Keung Tang

The rising importance of 3D understanding, pivotal in computer vision, autonomous driving, and robotics, is evident. However, a prevailing trend, which straightforwardly resorted to transferring 2D alignment strategies to the 3D domain,…

计算机视觉与模式识别 · 计算机科学 2024-01-26 Jiayi Ji , Haowei Wang , Changli Wu , Yiwei Ma , Xiaoshuai Sun , Rongrong Ji

Large Vision-Language Models (LVLMs) are an extension of Large Language Models (LLMs) that facilitate processing both image and text inputs, expanding AI capabilities. However, LVLMs struggle with object hallucinations due to their reliance…

计算与语言 · 计算机科学 2024-08-12 Avshalom Manevich , Reut Tsarfaty

Text-to-shape retrieval is an increasingly relevant problem with the growth of 3D shape data. Recent work on contrastive losses for learning joint embeddings over multimodal data has been successful at tasks such as retrieval and…

计算机视觉与模式识别 · 计算机科学 2023-12-29 Yue Ruan , Han-Hung Lee , Yiming Zhang , Ke Zhang , Angel X. Chang

Recent advancements in vision-language pre-training (e.g. CLIP) have shown that vision models can benefit from language supervision. While many models using language modality have achieved great success on 2D vision tasks, the joint…

计算机视觉与模式识别 · 计算机科学 2023-01-19 Rui Huang , Xuran Pan , Henry Zheng , Haojun Jiang , Zhifeng Xie , Shiji Song , Gao Huang

Vision-Language Instruction Tuning (VLIT) is a critical training phase for Large Vision-Language Models (LVLMs). With the improving capabilities of open-source LVLMs, researchers have increasingly turned to generate VLIT data by using…

计算机视觉与模式识别 · 计算机科学 2024-07-03 Ji Ma , Wei Suo , Peng Wang , Yanning Zhang

The task of LiDAR-based 3D Open-Vocabulary Detection (3D OVD) requires the detector to learn to detect novel objects from point clouds without off-the-shelf training labels. Previous methods focus on the learning of object-level…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Xingyu Peng , Si Liu , Chen Gao , Yan Bai , Beipeng Mu , Xiaofei Wang , Huaxia Xia

Multimodal 3D object detectors leverage the strengths of both geometry-aware LiDAR point clouds and semantically rich RGB images to enhance detection performance. However, the inherent heterogeneity between these modalities, including…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Zhuoqun Su , Huimin Lu , Shuaifeng Jiao , Junhao Xiao , Yaonan Wang , Xieyuanli Chen

Understanding 3D point clouds through language remains a fundamental challenge in computer graphics and visual computing, due to the irregular structure of point cloud data and the lack of explicit reasoning in existing 3D multimodal…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Chaoqi Chen , Qile Xu , Wenjun Zhou , Hui Huang

Large language models have demonstrated remarkable reasoning capabilities across diverse natural language tasks. However, comparable breakthroughs in scientific discovery are more limited, because understanding complex physical phenomena…

机器学习 · 计算机科学 2025-10-27 Jiyu Cui , Fang Wu , Haokai Zhao , Minggao Feng , Xenophon Evangelopoulos , Andrew I. Cooper , Yejin Choi