中文
相关论文

相关论文: VLMFusionOcc3D: VLM Assisted Multi-Modal 3D Semant…

200 篇论文

In autonomous driving, Vision Language Models (VLMs) excel at high-level reasoning , whereas semantic occupancy provides fine-grained details. Despite significant progress in individual fields, there is still no method that can effectively…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Chenxu Dang , Jie Wang , Guang Li , Zhiwen Hou , Zihan You , Hangjun Ye , Jie Ma , Long Chen , Yan Wang

3D object detection and occupancy prediction are critical tasks in autonomous driving, attracting significant attention. Despite the potential of recent vision-based methods, they encounter challenges under adverse conditions. Thus,…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Lianqing Zheng , Jianan Liu , Runwei Guan , Long Yang , Shouyi Lu , Yuanzhe Li , Xiaokai Bai , Jie Bai , Zhixiong Ma , Hui-Liang Shen , Xichan Zhu

3D semantic occupancy prediction has become a crucial perception task for comprehensive scene understanding in autonomous driving. While recent advances have explored 3D Gaussian splatting for occupancy modeling to substantially reduce…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Xiaoyang Yan , Muleilan Pei , Shaojie Shen

3D semantic occupancy prediction offers an intuitive and efficient scene understanding and has attracted significant interest in autonomous driving perception. Existing approaches either rely on full supervision, which demands costly…

计算机视觉与模式识别 · 计算机科学 2026-01-23 Naiyu Fang , Zheyuan Zhou , Fayao Liu , Xulei Yang , Jiacheng Wei , Lemiao Qiu , Hongsheng Li , Guosheng Lin

LiDAR-camera fusion can enhance the performance of 3D object detection by utilizing complementary information between depth-aware LiDAR points and semantically rich images. Existing voxel-based methods face significant challenges when…

计算机视觉与模式识别 · 计算机科学 2025-03-05 Ziying Song , Guoxin Zhang , Jun Xie , Lin Liu , Caiyan Jia , Shaoqing Xu , Zhepeng Wang

Occupancy prediction plays a pivotal role in autonomous driving (AD) due to the fine-grained geometric perception and general object recognition capabilities. However, existing methods often incur high computational costs, which contradicts…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Yulin He , Wei Chen , Tianci Xun , Yusong Tan

The lack of a large-scale 3D-text corpus has led recent works to distill open-vocabulary knowledge from vision-language models (VLMs). However, these methods typically rely on a single VLM to align the feature spaces of 3D models within a…

计算机视觉与模式识别 · 计算机科学 2025-03-31 Jinlong Li , Cristiano Saltori , Fabio Poiesi , Nicu Sebe

Vision-based 3D occupancy prediction is significantly challenged by the inherent limitations of monocular vision in depth estimation. This paper introduces CVT-Occ, a novel approach that leverages temporal fusion through the geometric…

计算机视觉与模式识别 · 计算机科学 2024-09-26 Zhangchen Ye , Tao Jiang , Chenfeng Xu , Yiming Li , Hang Zhao

Vision-based 3D semantic occupancy prediction is vital for autonomous driving, enabling unified modeling of static infrastructure and dynamic agents. Global occupancy maps serve as long-term memory priors, providing valuable historical…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Shanshuai Yuan , Julong Wei , Muer Tie , Xiangyun Ren , Zhongxue Gan , Wenchao Ding

3D semantic occupancy has rapidly become a research focus in the fields of robotics and autonomous driving environment perception due to its ability to provide more realistic geometric perception and its closer integration with downstream…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Mu Chen , Wenyu Chen , Mingchuan Yang , Yuan Zhang , Tao Han , Xinchi Li , Yunlong Li , Huaici Zhao

Multimodal 3D occupancy prediction has garnered significant attention for its potential in autonomous driving. However, most existing approaches are single-modality: camera-based methods lack depth information, while LiDAR-based methods…

计算机视觉与模式识别 · 计算机科学 2025-07-24 Zaipeng Duan , Chenxu Dang , Xuzhong Hu , Pei An , Junfeng Ding , Jie Zhan , Yunbiao Xu , Jie Ma

3D semantic occupancy and flow prediction are fundamental to spatiotemporal scene understanding. This paper proposes a vision-based framework with three targeted improvements. First, we introduce an occlusion-aware adaptive lifting…

计算机视觉与模式识别 · 计算机科学 2025-09-11 Dubing Chen , Jin Fang , Wencheng Han , Xinjing Cheng , Junbo Yin , Chenzhong Xu , Fahad Shahbaz Khan , Jianbing Shen

Vision-based 3D occupancy prediction has become a popular research task due to its versatility and affordability. Nowadays, conventional methods usually project the image-based vision features to 3D space and learn the geometric information…

计算机视觉与模式识别 · 计算机科学 2024-12-12 Yubo Cui , Zhiheng Li , Jiaqiang Wang , Zheng Fang

Driven by autonomous driving's demands for precise 3D perception, 3D semantic occupancy prediction has become a pivotal research topic. Unlike bird's-eye-view (BEV) methods, which restrict scene representation to a 2D plane, occupancy…

计算机视觉与模式识别 · 计算机科学 2025-09-05 Han Huang , Han Sun , Ningzhong Liu , Huiyu Zhou , Jiaquan Shen

With the growing number and diversity of Vision-Language Models (VLMs), many works explore language-based ensemble, collaboration, and routing techniques across multiple VLMs to improve multi-model reasoning. In contrast, we address the…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Selim Furkan Tekin , Yichang Xu , Gaowen Liu , Ramana Rao Kompella , Margaret L. Loper , Ling Liu

Human drivers adeptly navigate complex scenarios by utilizing rich attentional semantics, but the current autonomous systems struggle to replicate this ability, as they often lose critical semantic information when converting 2D…

计算机视觉与模式识别 · 计算机科学 2025-09-19 Pei Liu , Haipeng Liu , Haichao Liu , Xin Liu , Jinxin Ni , Jun Ma

Autonomous driving requires forecasting both geometry and semantics over time to effectively reason about future environment states. Existing vision-based occupancy forecasting methods focus on motion-related categories such as static and…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Riya Mohan , Juana Valeria Hurtado , Rohit Mohan , Abhinav Valada

Vision-Language-Action (VLA) models leverage Multimodal Large Language Models (MLLMs) for robotic control, but recent studies reveal that MLLMs exhibit limited spatial intelligence due to training predominantly on 2D data, resulting in…

Collaborative perception in automated vehicles leverages the exchange of information between agents, aiming to elevate perception results. Previous camera-based collaborative 3D perception methods typically employ 3D bounding boxes or…

计算机视觉与模式识别 · 计算机科学 2024-04-26 Rui Song , Chenwei Liang , Hu Cao , Zhiran Yan , Walter Zimmer , Markus Gross , Andreas Festag , Alois Knoll

Robot vision has greatly benefited from advancements in multimodal fusion techniques and vision-language models (VLMs). We adopt a task-oriented perspective to systematically review the applications and advancements of multimodal fusion…