中文
相关论文

相关论文: FreeOcc: Training-free Panoptic Occupancy Predicti…

200 篇论文

Modern 3D object detection datasets are constrained by narrow class taxonomies and costly manual annotations, limiting their ability to scale to open-world settings. In contrast, 2D vision-language models trained on web-scale image-text…

计算机视觉与模式识别 · 计算机科学 2025-07-21 Atharv Goel , Mehar Khurana

Understanding world dynamics is crucial for planning in autonomous driving. Recent methods attempt to achieve this by learning a 3D occupancy world model that forecasts future surrounding scenes based on current observation. However, 3D…

计算机视觉与模式识别 · 计算机科学 2025-02-12 Xiang Li , Pengfei Li , Yupeng Zheng , Wei Sun , Yan Wang , Yilun Chen

3D semantic occupancy prediction aims to forecast detailed geometric and semantic information of the surrounding environment for autonomous vehicles (AVs) using onboard surround-view cameras. Existing methods primarily focus on intricate…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Zhenxing Ming , Julie Stephany Berrio , Mao Shan , Stewart Worrall

Foundation models have exhibited unprecedented capabilities in tackling many domains and tasks. Models such as CLIP are currently widely used to bridge cross-modal representations, and text-to-image diffusion models are arguably the leading…

计算机视觉与模式识别 · 计算机科学 2025-11-19 Barbara Toniella Corradini , Mustafa Shukor , Paul Couairon , Guillaume Couairon , Franco Scarselli , Matthieu Cord

Closed-set 3D perception models trained on only a pre-defined set of object categories can be inadequate for safety critical applications such as autonomous driving where new object types can be encountered after deployment. In this paper,…

计算机视觉与模式识别 · 计算机科学 2023-09-27 Mahyar Najibi , Jingwei Ji , Yin Zhou , Charles R. Qi , Xinchen Yan , Scott Ettinger , Dragomir Anguelov

Detecting unknown objects in semantic segmentation is crucial for safety-critical applications such as autonomous driving. Large vision foundation models, including DINOv2, InternImage, and CLIP, have advanced visual representation learning…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Laith Nayal , Hadi Salloum , Ahmad Taha , Yaroslav Kholodov , Alexander Gasnikov

Multi-camera 3D perception has emerged as a prominent research field in autonomous driving, offering a viable and cost-effective alternative to LiDAR-based solutions. The existing multi-camera algorithms primarily rely on monocular 2D…

计算机视觉与模式识别 · 计算机科学 2024-04-30 Chen Min , Liang Xiao , Dawei Zhao , Yiming Nie , Bin Dai

3D semantic occupancy prediction aims to reconstruct the 3D geometry and semantics of the surrounding environment. With dense voxel labels, prior works typically formulate it as a dense segmentation task, independently classifying each…

图形学 · 计算机科学 2025-06-06 Wuyang Li , Zhu Yu , Alexandre Alahi

Embodied intelligence requires high-fidelity simulation environments to support perception and decision-making, yet existing platforms often suffer from data contamination and limited flexibility. To mitigate this, we propose…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Lechao Zhang , Haoran Xu , Jingyu Gong , Xuhong Wang , Yuan Xie , Xin Tan

For autonomous vehicles to proactively plan safe trajectories and make informed decisions, they must be able to predict the future occupancy states of the local environment. However, common issues with occupancy prediction include…

机器人学 · 计算机科学 2024-04-15 Maneekwan Toyungyernsub , Esen Yel , Jiachen Li , Mykel J. Kochenderfer

Autonomous navigation and exploration in unmapped environments remains a significant challenge in robotics due to the difficulty robots face in making commonsense inference of unobserved geometries. Recent advancements have demonstrated…

机器人学 · 计算机科学 2024-09-18 Alec Reed , Lorin Achey , Brendan Crowe , Bradley Hayes , Christoffer Heckman

In the field of autonomous driving, accurate and comprehensive perception of the 3D environment is crucial. Bird's Eye View (BEV) based methods have emerged as a promising solution for 3D object detection using multi-view images as input.…

计算机视觉与模式识别 · 计算机科学 2024-01-09 Qiu Zhou , Jinming Cao , Hanchao Leng , Yifang Yin , Yu Kun , Roger Zimmermann

Vision foundation models such as Contrastive Vision-Language Pre-training (CLIP) and Segment Anything (SAM) have demonstrated impressive zero-shot performance on image classification and segmentation tasks. However, the incorporation of…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Runnan Chen , Youquan Liu , Lingdong Kong , Nenglun Chen , Xinge Zhu , Yuexin Ma , Tongliang Liu , Wenping Wang

Controlling the spatial and semantic structure of diffusion-generated images remains a challenge. Existing methods like ControlNet rely on handcrafted condition maps and retraining, limiting flexibility and generalization. Inversion-based…

计算机视觉与模式识别 · 计算机科学 2025-11-10 Jiang Lin , Xinyu Chen , Song Wu , Zhiqiu Zhang , Jizhi Zhang , Ye Wang , Qiang Tang , Qian Wang , Jian Yang , Zili Yi

Understanding and reconstructing the 3D world through omnidirectional perception is an inevitable trend in the development of autonomous agents and embodied intelligence. However, existing 3D occupancy prediction methods are constrained by…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Mengfei Duan , Hao Shi , Fei Teng , Guoqiang Zhao , Yuheng Zhang , Zhiyong Li , Kailun Yang

Panoramic imagery provides holistic 360{\deg} visual coverage for perception in quadruped robots. However, existing occupancy prediction methods are mainly designed for wheeled autonomous driving and rely heavily on RGB cues, limiting their…

机器人学 · 计算机科学 2026-03-16 Guoqiang Zhao , Zhe Yang , Sheng Wu , Fei Teng , Mengfei Duan , Yuanfan Zheng , Kai Luo , Kailun Yang

Predicting variations in complex traffic environments is crucial for the safety of autonomous driving. Recent advancements in occupancy forecasting have enabled forecasting future 3D occupied status in driving environments by observing…

计算机视觉与模式识别 · 计算机科学 2025-02-24 Junliang Chen , Huaiyuan Xu , Yi Wang , Lap-Pui Chau

In robotic applications, a key requirement for safe and efficient motion planning is the ability to map obstacle-free space in unknown, cluttered 3D environments. However, commodity-grade RGB-D cameras commonly used for sensing fail to…

While 3D object bounding box (bbox) representation has been widely used in autonomous driving perception, it lacks the ability to capture the precise details of an object's intrinsic geometry. Recently, occupancy has emerged as a promising…

计算机视觉与模式识别 · 计算机科学 2024-12-09 Chaoda Zheng , Feng Wang , Naiyan Wang , Shuguang Cui , Zhen Li

Open-vocabulary semantic segmentation requires assigning pixel-level semantic labels while supporting an open and unrestricted set of categories. Training-free CLIP-based approaches preserve strong zero-shot generalization but typically…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Mohamad Zamini , Diksha Shukla