中文
相关论文

相关论文: From Filters to VLMs: Benchmarking Defogging Metho…

200 篇论文

Visual grounding is a common vision task that involves grounding descriptive sentences to the corresponding regions of an image. Most existing methods use independent image-text encoding and apply complex hand-crafted modules or…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Ming Dai , Lingfeng Yang , Yihao Xu , Zhenhua Feng , Wankou Yang

Vision Language Models (VLMs) have shown remarkable capabilities in multimodal understanding, yet their susceptibility to perturbations poses a significant threat to their reliability in real-world applications. Despite often being…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Jia Fu , Yongtao Wu , Yihang Chen , Kunyu Peng , Xiao Zhang , Volkan Cevher , Sepideh Pashami , Anders Holst

Modern optical flow methods make use of salient scene feature points detected and matched within the scene as a basis for sparse-to-dense optical flow estimation. Current feature detectors however either give sparse, non uniform point…

计算机视觉与模式识别 · 计算机科学 2019-05-21 Felix Stephenson , Toby Breckon , Ioannis Katramados

Traditional supervised methods for detecting AI-generated images depend on large, curated datasets for training and fail to generalize to novel, out-of-domain image generators. As an alternative, we explore pre-trained Vision-Language…

机器学习 · 计算机科学 2026-01-27 Zoher Kachwala , Danishjeet Singh , Danielle Yang , Filippo Menczer

Object detection from images captured by Unmanned Aerial Vehicles (UAVs) is becoming increasingly useful. Despite the great success of the generic object detection methods trained on ground-to-ground images, a huge performance drop is…

计算机视觉与模式识别 · 计算机科学 2020-10-06 Zhenyu Wu , Karthik Suresh , Priya Narayanan , Hongyu Xu , Heesung Kwon , Zhangyang Wang

While learning based compression techniques for images have outperformed traditional methods, they have not been widely adopted in machine learning pipelines. This is largely due to lack of standardization and lack of retention of salient…

图像与视频处理 · 电气工程与系统科学 2024-10-01 Kartik Gupta , Kimberley Faria , Vikas Mehta

Vision-language models (VLMs) have demonstrated impressive performance by effectively integrating visual and textual information to solve complex tasks. However, it is not clear how these models reason over the visual and textual data…

人工智能 · 计算机科学 2025-04-15 Pouya Pezeshkpour , Moin Aminnaseri , Estevam Hruschka

Autonomous vehicle navigation is a key challenge in artificial intelligence, requiring robust and accurate decision-making processes. This research introduces a new end-to-end method that exploits multimodal information from a single…

计算机视觉与模式识别 · 计算机科学 2024-09-20 Fouad Makiyeh , Mark Bastourous , Anass Bairouk , Wei Xiao , Mirjana Maras , Tsun-Hsuan Wangb , Marc Blanchon , Ramin Hasani , Patrick Chareyre , Daniela Rus

Many vision-language models (VLMs) that prove very effective at a range of multimodal task, build on CLIP-based vision encoders, which are known to have various limitations. We investigate the hypothesis that the strong language backbone in…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Sho Takishita , Jay Gala , Abdelrahman Mohamed , Kentaro Inui , Yova Kementchedjhieva

We present a method to improve the visual realism of low-quality, synthetic images, e.g. OpenGL renderings. Training an unpaired synthetic-to-real translation network in image space is severely under-constrained and produces visible…

计算机视觉与模式识别 · 计算机科学 2020-03-31 Sai Bi , Kalyan Sunkavalli , Federico Perazzi , Eli Shechtman , Vladimir Kim , Ravi Ramamoorthi

Accurate 6-DoF object pose estimation and tracking are critical for reliable robotic manipulation. However, zero-shot methods often fail under viewpoint-induced ambiguities and fixed-camera setups struggle when objects move or become…

机器人学 · 计算机科学 2026-03-10 Sheng Liu , Zhe Li , Weiheng Wang , Han Sun , Heng Zhang , Hongpeng Chen , Yusen Qin , Arash Ajoudani , Yizhao Wang

Detecting and evaluating surface coating defects is important for marine vessel maintenance. Currently, the assessment is carried out manually by qualified inspectors using international standards and their own experience. Automating the…

计算机视觉与模式识别 · 计算机科学 2022-03-21 Li Yu , Kareem Metwaly , James Z. Wang , Vishal Monga

Typically, object detection methods for autonomous driving that rely on supervised learning make the assumption of a consistent feature distribution between the training and testing data, this such assumption may fail in different weather…

计算机视觉与模式识别 · 计算机科学 2024-08-22 Jinlong Li , Runsheng Xu , Xinyu Liu , Jin Ma , Baolu Li , Qin Zou , Jiaqi Ma , Hongkai Yu

Effectively understanding urban scenes requires fine-grained spatial reasoning about objects, layouts, and depth cues. However, how well current vision-language models (VLMs), pretrained on general scenes, transfer these abilities to urban…

计算机视觉与模式识别 · 计算机科学 2025-09-01 Juneyoung Ro , Namwoo Kim , Yoonjin Yoon

As Vision-Language Models (VLMs) are increasingly deployed as autonomous cognitive cores for embodied assistants, evaluating their privacy awareness in physical environments becomes critical. Unlike digital chatbots, these agents operate in…

密码学与安全 · 计算机科学 2026-05-11 Junran Wang , Xinjie Shen , Zehao Jin , Pan Li

Despite tremendous advancements, current state-of-the-art Vision-Language Models (VLMs) are still far from perfect. They tend to hallucinate and may generate biased responses. In such circumstances, having a way to assess the reliability of…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Qian Yang , Weixiang Yan , Aishwarya Agrawal

Autonomous vehicles and driving systems use scene parsing as an essential tool to understand the surrounding environment. Panoptic segmentation is a state-of-the-art technique which proves to be pivotal in this use case. Deep learning-based…

计算机视觉与模式识别 · 计算机科学 2023-06-27 Ankur Chrungoo

Image contrast enhancement for outdoor vision is important for smart car auxiliary transport systems. The video frames captured in poor weather conditions are often characterized by poor visibility. Most image dehazing algorithms consider…

计算机视觉与模式识别 · 计算机科学 2015-10-06 Huimin Lu , Yujie Li , Shota Nakashima , Seiichi Serikawa

In this paper, we propose a novel framework for enhancing visual comprehension in autonomous driving systems by integrating visual language models (VLMs) with additional visual perception module specialised in object detection. We extend…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Linfeng He , Yiming Sun , Sihao Wu , Jiaxu Liu , Xiaowei Huang

Correlation Filters (CFs) have recently demonstrated excellent performance in terms of rapidly tracking objects under challenging photometric and geometric variations. The strength of the approach comes from its ability to efficiently learn…

计算机视觉与模式识别 · 计算机科学 2017-03-23 Hamed Kiani Galoogahi , Ashton Fagg , Simon Lucey