中文
相关论文

相关论文: ModalFormer: Multimodal Transformer for Low-Light …

200 篇论文

Matching visible and near-infrared (NIR) images remains a significant challenge in remote sensing image fusion. The nonlinear radiometric differences between heterogeneous remote sensing images make the image matching task even more…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Wang Zhang , Tingting Li , Yuntian Zhang , Gensheng Pei , Xiruo Jiang , Yazhou Yao

Most existing cross-modality person re-identification works rely on discriminative modality-shared features for reducing cross-modality variations and intra-modality variations. Despite some initial success, such modality-shared appearance…

计算机视觉与模式识别 · 计算机科学 2021-04-26 Nianchang Huang , Jianan Liu , Qiang Zhang , Jungong Han

This paper studies 3D low-dose computed tomography (CT) imaging. Although various deep learning methods were developed in this context, typically they focus on 2D images and perform denoising due to low-dose and deblurring for…

图像与视频处理 · 电气工程与系统科学 2024-01-10 Zhihao Chen , Chuang Niu , Qi Gao , Ge Wang , Hongming Shan

Due to the nature of enhancement--the absence of paired ground-truth information, high-level vision tasks have been recently employed to evaluate the performance of low-light image enhancement. A widely-used manner is to see how accurately…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Mingjia Li , Hao Zhao , Xiaojie Guo

In multimedia understanding tasks, corrupted samples pose a critical challenge, because when fed to machine learning models they lead to performance degradation. In the past, three groups of approaches have been proposed to handle noisy…

计算机视觉与模式识别 · 计算机科学 2024-03-01 Francesco Barbato , Umberto Michieli , Mehmet Kerim Yucel , Pietro Zanuttigh , Mete Ozay

Low-light images frequently occur due to unavoidable environmental influences or technical limitations, such as insufficient lighting or limited exposure time. To achieve better visibility for visual perception, low-light image enhancement…

计算机视觉与模式识别 · 计算机科学 2024-02-27 Shilv Cai , Liqun Chen , Sheng Zhong , Luxin Yan , Jiahuan Zhou , Xu Zou

This paper presents the first-ever study of adapting compressed image latents to suit the needs of downstream vision tasks that adopt Multimodal Large Language Models (MLLMs). MLLMs have extended the success of large language models to…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Chia-Hao Kao , Cheng Chien , Yu-Jen Tseng , Yi-Hsin Chen , Alessandro Gnutti , Shao-Yuan Lo , Wen-Hsiao Peng , Riccardo Leonardi

In reality, images often exhibit multiple degradations, such as rain and fog at night (triple degradations). However, in many cases, individuals may not want to remove all degradations, for instance, a blurry lens revealing a beautiful…

计算机视觉与模式识别 · 计算机科学 2024-04-17 Runwei Guan , Rongsheng Hu , Zhuhao Zhou , Tianlang Xue , Ka Lok Man , Jeremy Smith , Eng Gee Lim , Weiping Ding , Yutao Yue

Image restoration aims to reconstruct the latent clear images from their degraded versions. Despite the notable achievement, existing methods predominantly focus on handling specific degradation types and thus require specialized models,…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Xuanhua He , Lang Li , Yingying Wang , Hui Zheng , Ke Cao , Keyu Yan , Rui Li , Chengjun Xie , Jie Zhang , Man Zhou

Multimodal large language models (MLLMs) deployed on devices must adapt to continuously changing visual scenarios such as variations in background and perspective, to effectively perform complex visual tasks. To investigate catastrophic…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Kai Jiang , Siqi Huang , Xiangyu Chen , Jiawei Shao , Hongyuan Zhang , Ping Luo , Xuelong Li

Currently, most low-light image enhancement methods only consider information from a single view, neglecting the correlation between cross-view information. Therefore, the enhancement results produced by these methods are often…

计算机视觉与模式识别 · 计算机科学 2024-08-21 Linlin Hu , Ao Sun , Shijie Hao , Richang Hong , Meng Wang

Cross-modality magnetic resonance (MR) image synthesis can be used to generate missing modalities from given ones. Existing (supervised learning) methods often require a large number of paired multi-modal data to train an effective…

图像与视频处理 · 电气工程与系统科学 2023-06-21 Yonghao Li , Tao Zhou , Kelei He , Yi Zhou , Dinggang Shen

Recent advances in portable imaging have made camera-based screen capture ubiquitous. Unfortunately, frequency aliasing between the camera's color filter array (CFA) and the display's sub-pixels induces moir\'e patterns that severely…

计算机视觉与模式识别 · 计算机科学 2025-12-29 Jeahun Sung , Changhyun Roh , Chanho Eom , Jihyong Oh

Feature matching is a cornerstone task in computer vision, essential for applications such as image retrieval, stereo matching, 3D reconstruction, and SLAM. This survey comprehensively reviews modality-based feature matching, exploring…

计算机视觉与模式识别 · 计算机科学 2025-07-31 Weide Liu , Wei Zhou , Jun Liu , Ping Hu , Jun Cheng , Jungong Han , Weisi Lin

Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in multimodal tasks. Despite their impressive performance, MLLMs suffer from the modality imbalance issue, where visual information is often underutilized…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Hengzhuang Li , Xinsong Zhang , Qiming Peng , Bin Luo , Han Hu , Dengyang Jiang , Han-Jia Ye , Teng Zhang , Hai Jin

Low-light image enhancement (LLE) aims to improve the visual quality of images captured in poorly lit conditions, which often suffer from low brightness, low contrast, noise, and color distortions. These issues hinder the performance of…

计算机视觉与模式识别 · 计算机科学 2025-04-22 Junyu Xia , Jiesong Bai , Yihang Dong

Recent advances in semantic segmentation of multi-modal remote sensing images have significantly improved the accuracy of tree cover mapping, supporting applications in urban planning, forest monitoring, and ecological assessment.…

图像与视频处理 · 电气工程与系统科学 2025-12-16 Yuanyuan Gui , Wei Li , Yinjian Wang , Xiang-Gen Xia , Mauro Marty , Christian Ginzler , Zuyuan Wang

In order to get raw images of high quality for downstream Image Signal Process (ISP), in this paper we present an Efficient Locally Multiplicative Transformer called ELMformer for raw image restoration. ELMformer contains two core designs…

计算机视觉与模式识别 · 计算机科学 2022-09-01 Jiaqi Ma , Shengyuan Yan , Lefei Zhang , Guoli Wang , Qian Zhang

Cross-modal alignment Learning integrates information from different modalities like text, image, audio and video to create unified models. This approach develops shared representations and learns correlations between modalities, enabling…

计算机视觉与模式识别 · 计算机科学 2024-09-19 Bilal Faye , Hanane Azzag , Mustapha Lebbah

LiDAR data pretraining offers a promising approach to leveraging large-scale, readily available datasets for enhanced data utilization. However, existing methods predominantly focus on sparse voxel representation, overlooking the…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Xiang Xu , Lingdong Kong , Hui Shuai , Liang Pan , Ziwei Liu , Qingshan Liu