English
Related papers

Related papers: ModalFormer: Multimodal Transformer for Low-Light …

200 papers

In the Fourier frequency domain, luminance information is primarily encoded in the amplitude component, while spatial structure information is significantly contained within the phase component. Existing low-light image enhancement…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Tongshun Zhang , Pingping Liu , Ming Zhao , Haotian Lv

With the bloom of Large Language Models (LLMs), Multimodal Large Language Models (MLLMs) that incorporate LLMs with pre-trained vision models have recently demonstrated impressive performance across diverse vision-language tasks. However,…

Computation and Language · Computer Science 2026-01-13 Ziyue Wang , Chi Chen , Yiqi Zhu , Fuwen Luo , Peng Li , Ming Yan , Ji Zhang , Fei Huang , Maosong Sun , Yang Liu

Low-Light Image Enhancement (LLIE) is a crucial computer vision task that aims to restore detailed visual information from corrupted low-light images. Many existing LLIE methods are based on standard RGB (sRGB) space, which often produce…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Qingsen Yan , Yixu Feng , Cheng Zhang , Guansong Pang , Kangbiao Shi , Peng Wu , Wei Dong , Jinqiu Sun , Yanning Zhang

Despite the impressive advancements made in recent low-light image enhancement techniques, the scarcity of paired data has emerged as a significant obstacle to further advancements. This work proposes a mean-teacher-based semi-supervised…

Computer Vision and Pattern Recognition · Computer Science 2024-09-26 Guanlin Li , Ke Zhang , Ting Wang , Ming Li , Bin Zhao , Xuelong Li

Unlike single image task, stereo image enhancement can use another view information, and its key stage is how to perform cross-view feature interaction to extract useful information from another view. However, complex noise in low-light…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Minghua Zhao , Xiangdong Qin , Shuangli Du , Xuefei Bai , Jiahao Lyu , Yiguang Liu

In real-world clinical settings, magnetic resonance imaging (MRI) frequently suffers from missing modalities due to equipment variability or patient cooperation issues, which can significantly affect model performance. To address this…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Zhejia Zhang , Junjie Wang , Le Zhang

The event camera, benefiting from its high dynamic range and low latency, provides performance gain for low-light image enhancement. Unlike frame-based cameras, it records intensity changes with extremely high temporal resolution, capturing…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Chunyan She , Fujun Han , Chengyu Fang , Shukai Duan , Lidan Wang

The visibility of real-world images is often limited by both low-light and low-resolution, however, these issues are only addressed in the literature through Low-Light Enhancement (LLE) and Super- Resolution (SR) methods. Admittedly, a…

Image and Video Processing · Electrical Eng. & Systems 2024-03-01 Ziyu Yue , Jiaxin Gao , Sihan Xie , Yang Liu , Zhixun Su

Recently, Multimodal Large Language Models (MLLMs) have demonstrated impressive performance on instruction-following tasks by integrating pretrained visual encoders with large language models (LLMs). However, existing approaches often…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Wayner Barrios , Andrés Villa , Juan León Alcázar , SouYoung Jin , Bernard Ghanem

Low-dose CT (LDCT) images are often accompanied by significant noise, which negatively impacts image quality and subsequent diagnostic accuracy. To address the challenges of multi-scale feature fusion and diverse noise distribution patterns…

Image and Video Processing · Electrical Eng. & Systems 2025-05-20 Zhiting Zheng , Shuqi Wu , Wen Ding

Mixture of Vision Encoders (MoVE) has emerged as a powerful approach to enhance the fine-grained visual understanding of multimodal large language models (MLLMs), improving their ability to handle tasks such as complex optical character…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Mozhgan Nasr Azadani , James Riddell , Sean Sedwards , Krzysztof Czarnecki

Low-light image enhancement, such as recovering color and texture details from low-light images, is a complex and vital task. For automated driving, low-light scenarios will have serious implications for vision-based applications. To…

Image and Video Processing · Electrical Eng. & Systems 2021-09-01 Yangyang Qu , Kai Chen , Chao Liu , Yongsheng Ou

Multimodal Large Language Models (MLLMs) demonstrate remarkable image-language capabilities, but their widespread use faces challenges in cost-effective training and adaptation. Existing approaches often necessitate expensive language model…

Computer Vision and Pattern Recognition · Computer Science 2024-08-14 Sayna Ebrahimi , Sercan O. Arik , Tejas Nama , Tomas Pfister

Low-light images suffer from complex degradation, and existing enhancement methods often encode all degradation factors within a single latent space. This leads to highly entangled features and strong black-box characteristics, making the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-17 Shuangli Du , Siming Yan , Zhenghao Shi , Zhenzhen You , Lu Sun

Low-cost cross-modal representation learning is crucial for deriving semantic representations across diverse modalities such as text, audio, images, and video. Traditional approaches typically depend on large specialized models trained from…

Machine Learning · Computer Science 2024-09-10 Bilal Faye , Hanane Azzag , Mustapha Lebbah , Djamel Bouchaffra

Large Vision-Language Models (LVLMs) often omit or misrepresent critical visual content in generated image captions. Minimizing such information loss will force LVLMs to focus on image details to generate precise descriptions. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Haonan Jia , Shichao Dong , Xin Dong , Zenghui Sun , Jin Wang , Jinsong Lan , Xiaoyong Zhu , Bo Zheng , Kaifu Zhang

Low-Light Video Enhancement (LLVE) seeks to restore dynamic or static scenes plagued by severe invisibility and noise. In this paper, we present an innovative video decomposition strategy that incorporates view-independent and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Xiaogang Xu , Kun Zhou , Tao Hu , Jiafei Wu , Ruixing Wang , Hao Peng , Bei Yu

Multi-modal skin lesion diagnosis (MSLD) has achieved remarkable success by modern computer-aided diagnosis (CAD) technology based on deep convolutions. However, the information aggregation across modalities in MSLD remains challenging due…

Computer Vision and Pattern Recognition · Computer Science 2023-03-03 Yilan Zhang , Fengying Xie , Jianqi Chen

This paper presents a pure transformer-based approach, dubbed the Multi-Modal Video Transformer (MM-ViT), for video action recognition. Different from other schemes which solely utilize the decoded RGB frames, MM-ViT operates exclusively in…

Computer Vision and Pattern Recognition · Computer Science 2021-11-16 Jiawei Chen , Chiu Man Ho

Image feature matching, a foundational task in computer vision, remains challenging for multimodal image applications, often necessitating intricate training on specific datasets. In this paper, we introduce a Unified Feature Matching…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Yide Di , Yun Liao , Hao Zhou , Kaijun Zhu , Qing Duan , Junhui Liu , Mingyu Lu
‹ Prev 1 4 5 6 7 8 10 Next ›