中文
相关论文

相关论文: From Filters to VLMs: Benchmarking Defogging Metho…

200 篇论文

Understanding foggy image sequence in the driving scenes is critical for autonomous driving, but it remains a challenging task due to the difficulty in collecting and annotating real-world images of adverse weather. Recently, the…

计算机视觉与模式识别 · 计算机科学 2022-06-13 Liang Liao , Wenyi Chen , Jing Xiao , Zheng Wang , Chia-Wen Lin , Shin'ichi Satoh

Beyond the commonly recognized optical aberrations, the imaging performance of simplified optical systems--including single-lens and metalens designs--is often further degraded by veiling glare caused by stray-light scattering from…

图像与视频处理 · 电气工程与系统科学 2026-03-09 Xiaolong Qian , Qi Jiang , Lei Sun , Zongxi Yu , Kailun Yang , Peixuan Wu , Jiacheng Zhou , Yao Gao , Yaoguang Ma , Ming-Hsuan Yang , Kaiwei Wang

Vision language models (VLMs) are designed to extract relevant visuospatial information from images. Some research suggests that VLMs can exhibit humanlike scene understanding, while other investigations reveal difficulties in their ability…

The field of object detection and understanding is rapidly evolving, driven by advances in both traditional CNN-based models and emerging multi-modal large language models (LLMs). While CNNs like ResNet and YOLO remain highly effective for…

计算机视觉与模式识别 · 计算机科学 2025-10-13 Nirmal Elamon , Rouzbeh Davoudi

In real-world environments, outdoor imaging systems are often affected by disturbances such as rain degradation. Especially, in nighttime driving scenes, insufficient and uneven lighting shrouds the scenes in darkness, resulting degradation…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Cidan Shi , Lihuang Fang , Han Wu , Xiaoyu Xian , Yukai Shi , Liang Lin

Multimodal sensor fusion is an essential capability for autonomous robots, enabling object detection and decision-making in the presence of failing or uncertain inputs. While recent fusion methods excel in normal environmental conditions,…

计算机视觉与模式识别 · 计算机科学 2025-08-25 Edoardo Palladin , Roland Dietze , Praveen Narayanan , Mario Bijelic , Felix Heide

Robots with internal visual self-models promise unprecedented adaptability, yet existing autonomous modeling pipelines remain fragile under realistic sensing conditions such as noisy imagery and cluttered backgrounds. This paper presents…

机器人学 · 计算机科学 2025-10-07 Salim Rezvani , Ammar Jaleel Mahmood , Robin Chhabra

Recent advancements in perception for autonomous driving are driven by deep learning. In order to achieve robust and accurate scene understanding, autonomous vehicles are usually equipped with different sensors (e.g. cameras, LiDARs,…

Vision-language models like CLIP excel at recognizing the single, prominent object in a scene. However, they struggle in complex scenes containing multiple objects. We identify a fundamental reason for this limitation: VLM feature space…

计算机视觉与模式识别 · 计算机科学 2025-09-26 Samyak Rawlekar , Yujun Cai , Yiwei Wang , Ming-Hsuan Yang , Narendra Ahuja

The growing demand of industrial, automotive and service robots presents a challenge to the centralized Cloud Robotics model in terms of privacy, security, latency, bandwidth, and reliability. In this paper, we present a `Fog Robotics'…

机器人学 · 计算机科学 2019-03-25 Ajay Kumar Tanwani , Nitesh Mor , John Kubiatowicz , Joseph E. Gonzalez , Ken Goldberg

Visualization authoring is an iterative process requiring users to adjust parameters to achieve desired aesthetics. Due to its complexity, users often create defective visualizations and struggle to fix them. Many seek help on forums (e.g.,…

人机交互 · 计算机科学 2026-02-05 Shuyu Shen , Sirong Lu , Leixian Shen , Yuyu Luo

Video restoration for noise removal, deblurring or super-resolution is attracting more and more attention in the fields of image processing and computer vision. Works on video restoration with data-driven approaches for fog removal are rare…

计算机视觉与模式识别 · 计算机科学 2023-10-03 Alexandra Duminil , Jean-Philippe Tarel , Roland Brémond

Accurate classification of weather conditions in images is essential for enhancing the performance of object detection and classification models under varying weather conditions. This paper presents a comprehensive study on classifying…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Eden Ship , Eitan Spivak , Shubham Agarwal , Raz Birman , Ofer Hadar

Video camouflaged object detection (VCOD) is challenging due to dynamic environments. Existing methods face two main issues: (1) SAM-based methods struggle to separate camouflaged object edges due to model freezing, and (2) MLLM-based…

计算机视觉与模式识别 · 计算机科学 2025-09-09 Hua Zhang , Changjiang Luo , Ruoyu Chen

Large language models (LLMs) are growingly extended to process multimodal data such as text and video simultaneously. Their remarkable performance in understanding what is shown in images is surpassing specialized neural networks (NNs) such…

计算机视觉与模式识别 · 计算机科学 2025-07-24 Malsha Ashani Mahawatta Dona , Beatriz Cabrero-Daniel , Yinan Yu , Christian Berger

Vision Large Language Models (VLMs) combine visual understanding with natural language processing, enabling tasks like image captioning, visual question answering, and video analysis. While VLMs show impressive capabilities across domains…

计算机视觉与模式识别 · 计算机科学 2025-06-18 Ahmed Sharshar , Latif U. Khan , Waseem Ullah , Mohsen Guizani

Fluorescence microscopy is a major driver of scientific progress in the life sciences. Although high-end confocal microscopes are capable of filtering out-of-focus light, cheaper and more accessible microscopy modalities, such as widefield…

图像与视频处理 · 电气工程与系统科学 2026-04-08 Anirban Ray , Ashesh Ashesh , Florian Jug

The contemporary phenomenon of deepfakes, utilizing GAN or diffusion models for face swapping, presents a substantial and evolving threat in digital media, identity verification, and a multitude of other systems. The majority of existing…

计算机视觉与模式识别 · 计算机科学 2025-07-31 Viacheslav Pirogov

Images acquired by computer vision systems under low light conditions have multiple characteristics like high noise, lousy illumination, reflectance, and bad contrast, which make object detection tasks difficult. Much work has been done to…

计算机视觉与模式识别 · 计算机科学 2021-08-02 Winston Chen , Tejas Shah

Vision-language models (VLMs) often struggle to generate accurate and detailed captions for high-resolution images since they are typically pre-trained on low-resolution inputs (e.g., 224x224 or 336x336 pixels). Downscaling high-resolution…

计算机视觉与模式识别 · 计算机科学 2025-11-03 Hankyeol Lee , Gawon Seo , Kyounggyu Lee , Dogun Kim , Kyungwoo Song , Jiyoung Jung