中文
相关论文

相关论文: Deflickering Vision-Based Occupancy Networks throu…

200 篇论文

Vision-language navigation (VLN) requires an agent to navigate through an 3D environment based on visual observations and natural language instructions. It is clear that the pivotal factor for successful navigation lies in the comprehensive…

计算机视觉与模式识别 · 计算机科学 2024-03-22 Rui Liu , Wenguan Wang , Yi Yang

3D occupancy prediction based on multi-sensor fusion,crucial for a reliable autonomous driving system, enables fine-grained understanding of 3D scenes. Previous fusion-based 3D occupancy predictions relied on depth estimation for processing…

计算机视觉与模式识别 · 计算机科学 2024-07-11 Ji Zhang , Yiran Ding , Zixin Liu

Correlation filters take advantage of specific properties in the Fourier domain allowing them to be estimated efficiently: O(NDlogD) in the frequency domain, versus O(D^3 + ND^2) spatially where D is signal length, and N is the number of…

计算机视觉与模式识别 · 计算机科学 2014-04-01 Hamed Kiani Galoogahi , Terence Sim , Simon Lucey

Visual sensor networks are used for monitoring traffic in large cities and are promised to support automated driving in complex road segments. The pose of these sensors, i.e. position and orientation, directly determines the coverage of the…

计算机视觉与模式识别 · 计算机科学 2024-10-28 Eduardo Arnold , Sajjad Mozaffari , Mehrdad Dianati , Paul Jennings

Vision-based 3D occupancy prediction has become a popular research task due to its versatility and affordability. Nowadays, conventional methods usually project the image-based vision features to 3D space and learn the geometric information…

计算机视觉与模式识别 · 计算机科学 2024-12-12 Yubo Cui , Zhiheng Li , Jiaqiang Wang , Zheng Fang

Vision-language models (VLMs) frequently generate hallucinated content plausible but incorrect claims about image content. We propose a training-free self-correction framework enabling VLMs to iteratively refine responses through…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Kassoum Sanogo , Renzo Ardiccioni

We propose a fixed-lag smoother-based sensor fusion architecture to leverage the complementary benefits of range-based sensors and visual-inertial odometry (VIO) for localization. We use two fixed-lag smoothers (FLS) to decouple accurate…

机器人学 · 计算机科学 2024-01-05 Abhishek Goudar , Wenda Zhao , Angela P. Schoellig

In Vision Language Models (VLMs), vision tokens are quantity-heavy yet information-dispersed compared with language tokens, thus consume too much unnecessary computation. Pruning redundant vision tokens for high VLM inference efficiency has…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Guangyuan Li , Rongzhen Zhao , Jinhong Deng , Yanbo Wang , Joni Pajarinen

Trajectory prediction plays an important role in various applications, including autonomous driving, robotics, and scene understanding. Existing approaches mainly focus on developing compact neural networks to increase prediction precision…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Yi Xu , Yun Fu

Multi-sensor fusion significantly enhances the accuracy and robustness of 3D semantic occupancy prediction, which is crucial for autonomous driving and robotics. However, most existing approaches depend on high-resolution images and complex…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Zhen Yang , Yanpeng Dong , Jiayu Wang , Heng Wang , Lichao Ma , Zijian Cui , Qi Liu , Haoran Pei , Kexin Zhang , Chao Zhang

Video Diffusion Models (VDMs) have demonstrated remarkable capabilities in synthesizing realistic videos by learning from large-scale data. Although vanilla Low-Rank Adaptation (LoRA) can learn specific spatial or temporal movement to…

计算机视觉与模式识别 · 计算机科学 2025-07-09 Yisu Zhang , Chenjie Cao , Chaohui Yu , Jianke Zhu

Large Vision-Language Models (LVLMs) excel at captioning, visual question answering, and robotics by combining vision and language, yet they often miss obvious objects or hallucinate nonexistent ones in atypical scenes. We examine these…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Zhaoyang Li , Zhan Ling , Yuchen Zhou , Litian Gong , Erdem Bıyık , Hao Su

Recent large-scale Vision Language Action (VLA) models have shown superior performance in robotic manipulation tasks guided by natural language. However, current VLA models suffer from two drawbacks: (i) generation of massive tokens leading…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Juyi Lin , Amir Taherin , Arash Akbari , Arman Akbari , Lei Lu , Guangyu Chen , Taskin Padir , Xiaomeng Yang , Weiwei Chen , Yiqian Li , Xue Lin , David Kaeli , Pu Zhao , Yanzhi Wang

Visual place recognition methods struggle with occlusions and partial visual overlaps. We propose a novel visual place recognition approach based on overlap prediction, called VOP, shifting from traditional reliance on global image…

计算机视觉与模式识别 · 计算机科学 2024-12-05 Tong Wei , Philipp Lindenberger , Jiri Matas , Daniel Barath

The vision-based perception for autonomous driving has undergone a transformation from the bird-eye-view (BEV) representations to the 3D semantic occupancy. Compared with the BEV planes, the 3D semantic occupancy further provides structural…

计算机视觉与模式识别 · 计算机科学 2023-04-12 Yunpeng Zhang , Zheng Zhu , Dalong Du

Oblique images are aerial photographs taken at oblique angles to the earth's surface. Projections of vector and other geospatial data in these images depend on camera parameters, positions of the geospatial entities, surface terrain,…

计算机视觉与模式识别 · 计算机科学 2022-06-22 Pragyana Mishra , Eyal Ofek , Gur Kimchi

State-of-the-art neural network models estimate large displacement optical flow in multi-resolution and use warping to propagate the estimation between two resolutions. Despite their impressive results, it is known that there are two…

计算机视觉与模式识别 · 计算机科学 2019-03-05 Yao Lu , Jack Valmadre , Heng Wang , Juho Kannala , Mehrtash Harandi , Philip H. S. Torr

Dynamic scattering remains a significant challenge to the practical deployment of anti-scattering imaging. Existing methods, such as transmission matrix measurements, iterative wavefront shaping, and optical phase conjugation, depend on a…

LiDAR-based 3D object detectors have achieved unprecedented speed and accuracy in autonomous driving applications. However, similar to other neural networks, they are often biased toward high-confidence predictions or return detections…

计算机视觉与模式识别 · 计算机科学 2023-10-03 Aldi Piroli , Vinzenz Dallabetta , Johannes Kopp , Marc Walessa , Daniel Meissner , Klaus Dietmayer

Recent video semantic segmentation (VSS) methods have demonstrated promising results in well-lit environments. However, their performance significantly drops in low-light scenarios due to limited visibility and reduced contextual details.…

计算机视觉与模式识别 · 计算机科学 2026-01-29 Zhen Yao , Mooi Choo Chuah