English
Related papers

Related papers: Deflickering Vision-Based Occupancy Networks throu…

200 papers

In recent years, autonomous driving has garnered escalating attention for its potential to relieve drivers' burdens and improve driving safety. Vision-based 3D occupancy prediction, which predicts the spatial occupancy status and semantics…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Yanan Zhang , Jinqing Zhang , Zengran Wang , Junhao Xu , Di Huang

Vision-centric occupancy networks, which represent the surrounding environment with uniform voxels with semantics, have become a new trend for safe driving of camera-only autonomous driving perception systems, as they are able to detect…

Computer Vision and Pattern Recognition · Computer Science 2024-09-19 Yining Shi , Jiusi Li , Kun Jiang , Ke Wang , Yunlong Wang , Mengmeng Yang , Diange Yang

The safe operation of autonomous vehicles (AVs) is highly dependent on their understanding of the surroundings. For this, the task of 3D semantic occupancy prediction divides the space around the sensors into voxels, and labels each voxel…

Computer Vision and Pattern Recognition · Computer Science 2025-05-07 Zhenxing Ming , Julie Stephany Berrio , Mao Shan , Yaoqi Huang , Hongyu Lyu , Nguyen Hoang Khoi Tran , Tzu-Yun Tseng , Stewart Worrall

Crucial for autonomous exploration, online 3D occupancy prediction and mapping incrementally constructs dense spatial representations on the fly. However, recent Gaussian-centric methods struggle with structural boundary fidelity and rely…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Ruoyu Wang , Yong Liu , Sheng Tao , Yuhang Lin , Yukai Ma

The 3D occupancy estimation task has become an important challenge in the area of vision-based autonomous driving recently. However, most existing camera-based methods rely on costly 3D voxel labels or LiDAR scans for training, limiting…

Computer Vision and Pattern Recognition · Computer Science 2024-07-26 Simon Boeder , Fabian Gigengack , Benjamin Risse

Vision-and-Language Navigation (VLN) requires agents to follow long-horizon instructions and navigate complex 3D environments. However, existing approaches face two major challenges: constructing an effective long-term memory bank and…

Robotics · Computer Science 2026-03-27 Zihao Xin , Wentong Li , Yixuan Jiang , Bin Wang , Runmin Cong , Jie Qin , Shengjun Huang

Virtual and augmented reality (VR/AR) displays strive to provide a resolution, framerate and field of view that matches the perceptual capabilities of the human visual system, all while constrained by limited compute budgets and…

Human-Computer Interaction · Computer Science 2021-06-22 Brooke Krajancich , Petr Kellnhofer , Gordon Wetzstein

Self-supervised 3D occupancy prediction offers a promising solution for understanding complex driving scenes without requiring costly 3D annotations. However, training dense occupancy decoders to capture fine-grained geometry and semantics…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Fengyi Zhang , Xiangyu Sun , Huitong Yang , Zheng Zhang , Zi Huang , Yadan Luo

Optical neural networks offer a route to low-latency and energy-efficient inference by encoding computation in light propagation. However, most existing implementations rely on planar photonic circuits or discretely spaced diffractive…

An emerging paradigm in vision-and-language navigation (VLN) is the use of history-aware multi-modal transformer models. Given a language instruction, these models process observation and navigation history to predict the most appropriate…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Dongwoo Kang , Akhil Perincherry , Zachary Coalson , Aiden Gabriel , Stefan Lee , Sanghyun Hong

Recent progress in deep generative models has led to tremendous breakthroughs in image generation. However, while existing models can synthesize photorealistic images, they lack an understanding of our underlying 3D world. We present a new…

Computer Vision and Pattern Recognition · Computer Science 2018-12-07 Jun-Yan Zhu , Zhoutong Zhang , Chengkai Zhang , Jiajun Wu , Antonio Torralba , Joshua B. Tenenbaum , William T. Freeman

A comprehensive understanding of 3D scenes is crucial in autonomous vehicles (AVs), and recent models for 3D semantic occupancy prediction have successfully addressed the challenge of describing real-world objects with varied shapes and…

Computer Vision and Pattern Recognition · Computer Science 2024-05-10 Zhenxing Ming , Julie Stephany Berrio , Mao Shan , Stewart Worrall

Accurate perception of the dynamic environment is a fundamental task for autonomous driving and robot systems. This paper introduces Let Occ Flow, the first self-supervised work for joint 3D occupancy and occupancy flow prediction using…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Yili Liu , Linzhan Mou , Xuan Yu , Chenrui Han , Sitong Mao , Rong Xiong , Yue Wang

While multi-modal 3D semantic occupancy prediction typically enhances robustness by fusing camera and LiDAR inputs, its effectiveness is fundamentally constrained by environmental variability. Specifically, camera sensors suffer from severe…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 A. Enes Doruk , Abdelaziz Hussein , Hasan F. Ates

Vision-based perception for autonomous driving requires an explicit modeling of a 3D space, where 2D latent representations are mapped and subsequent 3D operators are applied. However, operating on dense latent spaces introduces a cubic…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Pin Tang , Zhongdao Wang , Guoqing Wang , Jilai Zheng , Xiangxuan Ren , Bailan Feng , Chao Ma

Recent progress in blind face restoration has resulted in producing high-quality restored results for static images. However, efforts to extend these advancements to video scenarios have been minimal, partly because of the absence of…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Zhouxia Wang , Jiawei Zhang , Xintao Wang , Tianshui Chen , Ying Shan , Wenping Wang , Ping Luo

Multimodal large language models (MLLMs) have shown strong vision-language reasoning abilities but still lack robust 3D spatial understanding, which is critical for autonomous driving. This limitation stems from two key challenges: (1) the…

Artificial Intelligence · Computer Science 2025-09-09 Ruixun Liu , Lingyu Kong , Derun Li , Hang Zhao

3D occupancy becomes a promising perception representation for autonomous driving to model the surrounding environment at a fine-grained scale. However, it remains challenging to efficiently aggregate 3D occupancy over time across multiple…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Ziyang Leng , Jiawei Yang , Wenlong Yi , Bolei Zhou

Data-driven visual odometry (VO) is a critical subroutine for autonomous edge robotics, and recent progress in the field has produced highly accurate point predictions in complex environments. However, emerging autonomous edge robotics…

Computer Vision and Pattern Recognition · Computer Science 2023-03-07 Alex C. Stutts , Danilo Erricolo , Theja Tulabandhula , Amit Ranjan Trivedi

Vision-based occupancy prediction, also known as 3D Semantic Scene Completion (SSC), presents a significant challenge in computer vision. Previous methods, confined to onboard processing, struggle with simultaneous geometric and semantic…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Hao Shi , Song Wang , Jiaming Zhang , Xiaoting Yin , Guangming Wang , Jianke Zhu , Kailun Yang , Kaiwei Wang