中文
相关论文

相关论文: VLMFusionOcc3D: VLM Assisted Multi-Modal 3D Semant…

200 篇论文

Crucial for autonomous exploration, online 3D occupancy prediction and mapping incrementally constructs dense spatial representations on the fly. However, recent Gaussian-centric methods struggle with structural boundary fidelity and rely…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Ruoyu Wang , Yong Liu , Sheng Tao , Yuhang Lin , Yukai Ma

Robust 3D occupancy prediction is essential for autonomous driving, particularly under adverse weather conditions where traditional vision-only systems struggle. While the fusion of surround-view 4D radar and cameras offers a promising…

计算机视觉与模式识别 · 计算机科学 2025-08-11 Long Yang , Lianqing Zheng , Wenjin Ai , Minghao Liu , Sen Li , Qunshu Lin , Shengyu Yan , Jie Bai , Zhixiong Ma , Tao Huang , Xichan Zhu

3D occupancy prediction has become a key perception task in autonomous driving, as it enables comprehensive scene understanding. Recent methods enhance this understanding by incorporating spatiotemporal information through multi-frame…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Seokha Moon , Janghyun Baek , Giseop Kim , Jinkyu Kim , Sunwook Choi

Large Vision Language Models (LVLMs) have shown strong capabilities in understanding and analyzing visual scenes across various domains. However, in the context of autonomous driving, their limited comprehension of 3D environments restricts…

计算机视觉与模式识别 · 计算机科学 2025-05-02 Jannik Lübberstedt , Esteban Rivera , Nico Uhlemann , Markus Lienkamp

Open-world 3D semantic occupancy prediction aims to generate a voxelized 3D representation from sensor inputs while recognizing both known and unknown objects. Transferring open-vocabulary knowledge from vision-language models (VLMs) offers…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Peizheng Li , Shuxiao Ding , You Zhou , Qingwen Zhang , Onat Inak , Larissa Triess , Niklas Hanselmann , Marius Cordts , Andreas Zell

Multi-sensor fusion significantly enhances the accuracy and robustness of 3D semantic occupancy prediction, which is crucial for autonomous driving and robotics. However, most existing approaches depend on high-resolution images and complex…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Zhen Yang , Yanpeng Dong , Jiayu Wang , Heng Wang , Lichao Ma , Zijian Cui , Qi Liu , Haoran Pei , Kexin Zhang , Chao Zhang

Vehicle-to-everything (V2X) cooperation has emerged as a promising paradigm to overcome the perception limitations of classical autonomous driving by leveraging information from both ego-vehicle and infrastructure sensors. However,…

机器人学 · 计算机科学 2025-06-23 Junwei You , Haotian Shi , Zhuoyu Jiang , Zilin Huang , Rui Gan , Keshu Wu , Xi Cheng , Xiaopeng Li , Bin Ran

The fusion of multimodal sensor data streams such as camera images and lidar point clouds plays an important role in the operation of autonomous vehicles (AVs). Robust perception across a range of adverse weather and lighting conditions is…

计算机视觉与模式识别 · 计算机科学 2024-05-27 Shounak Sural , Nishad Sahu , Ragunathan Rajkumar

In recent years, autonomous driving has garnered escalating attention for its potential to relieve drivers' burdens and improve driving safety. Vision-based 3D occupancy prediction, which predicts the spatial occupancy status and semantics…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Yanan Zhang , Jinqing Zhang , Zengran Wang , Junhao Xu , Di Huang

Robotic perception requires the modeling of both 3D geometry and semantics. Existing methods typically focus on estimating 3D bounding boxes, neglecting finer geometric details and struggling to handle general, out-of-vocabulary objects. 3D…

计算机视觉与模式识别 · 计算机科学 2023-12-14 Xiaoyu Tian , Tao Jiang , Longfei Yun , Yucheng Mao , Huitong Yang , Yue Wang , Yilun Wang , Hang Zhao

The performance of multi-modal 3D occupancy prediction is limited by ineffective fusion, mainly due to geometry-semantics mismatch from fixed fusion strategies and surface detail loss caused by sparse, noisy annotations. The mismatch stems…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Luyao Lei , Shuo Xu , Yifan Bai , Xing Wei

Visual Odometry (VO) plays a pivotal role in autonomous systems, with a principal challenge being the lack of depth information in camera images. This paper introduces OCC-VO, a novel framework that capitalizes on recent advances in deep…

机器人学 · 计算机科学 2024-03-27 Heng Li , Yifan Duan , Xinran Zhang , Haiyi Liu , Jianmin Ji , Yanyong Zhang

Semantic segmentation in autonomous driving has been undergoing an evolution from sparse point segmentation to dense voxel segmentation, where the objective is to predict the semantic occupancy of each voxel in the concerned 3D space. The…

计算机视觉与模式识别 · 计算机科学 2023-09-01 Sicheng Zuo , Wenzhao Zheng , Yuanhui Huang , Jie Zhou , Jiwen Lu

Vision-language models (VLMs) have shown promise in 2D medical image analysis, but extending them to 3D remains challenging due to the high computational demands of volumetric data and the difficulty of aligning 3D spatial features with…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Yu Xin , Gorkem Can Ates , Kuang Gong , Wei Shao

Panoramic imagery provides holistic 360{\deg} visual coverage for perception in quadruped robots. However, existing occupancy prediction methods are mainly designed for wheeled autonomous driving and rely heavily on RGB cues, limiting their…

机器人学 · 计算机科学 2026-03-16 Guoqiang Zhao , Zhe Yang , Sheng Wu , Fei Teng , Mengfei Duan , Yuanfan Zheng , Kai Luo , Kailun Yang

Camera-based 3D semantic occupancy prediction offers an efficient and cost-effective solution for perceiving surrounding scenes in autonomous driving. However, existing works rely on explicit occupancy state inference, leading to numerous…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Naiyu Fang , Zheyuan Zhou , Kang Wang , Ruibo Li , Lemiao Qiu , Shuyou Zhang , Zhe Wang , Guosheng Lin

The prediction of 3D semantic occupancy enables autonomous vehicles (AVs) to perceive the fine-grained geometric and semantic scene structure for safe navigation and decision-making. Existing methods mainly rely on either voxel-based…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Zhenxing Ming , Yaoqi Huang , Julie Stephany Berrio , Mao Shan , Stewart Worrall

The rise of multi-modal large language models(MLLMs) has spurred their applications in autonomous driving. Recent MLLM-based methods perform action by learning a direct mapping from perception to action, neglecting the dynamics of the world…

计算机视觉与模式识别 · 计算机科学 2024-09-06 Julong Wei , Shanshuai Yuan , Pengfei Li , Qingda Hu , Zhongxue Gan , Wenchao Ding

Multi-view cooperative perception and multimodal fusion are essential for reliable 3D spatiotemporal understanding in autonomous driving, especially under occlusions, limited viewpoints, and communication delays in V2X scenarios. This paper…

计算机视觉与模式识别 · 计算机科学 2025-12-29 Zhenwei Yang , Yibo Ai , Weidong Zhang

3D semantic occupancy prediction is an essential part of autonomous driving, focusing on capturing the geometric details of scenes. Off-road environments are rich in geometric information, therefore it is suitable for 3D semantic occupancy…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Heng Zhai , Jilin Mei , Chen Min , Liang Chen , Fangzhou Zhao , Yu Hu