English
Related papers

Related papers: iDETEX: Empowering MLLMs for Intelligent DETailed …

200 papers

Image quality scoring and interpreting are two fundamental components of Image Quality Assessment (IQA). The former quantifies image quality, while the latter enables descriptive question answering about image quality. Traditionally, these…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Zicheng Zhang , Haoning Wu , Ziheng Jia , Weisi Lin , Guangtao Zhai

Image Quality Assessment (IQA) and Image Aesthetic Assessment (IAA) aim to simulate human subjective perception of image visual quality and aesthetic appeal. Despite distinct learning objectives, they have underlying interconnectedness due…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Hantao Zhou , Longxiang Tang , Rui Yang , Guanyi Qin , Yan Zhang , Yutao Li , Xiu Li , Runze Hu , Guangtao Zhai

Image quality assessment (IQA) focuses on the perceptual visual quality of images, playing a crucial role in downstream tasks such as image reconstruction, compression, and generation. The rapid advancement of multi-modal large language…

Computer Vision and Pattern Recognition · Computer Science 2025-05-26 Weiqi Li , Xuanyu Zhang , Shijie Zhao , Yabin Zhang , Junlin Li , Li Zhang , Jian Zhang

In this paper, we introduce ILLUME, a unified multimodal large language model (MLLM) that seamlessly integrates multimodal understanding and generation capabilities within a single large language model through a unified next-token…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Chunwei Wang , Guansong Lu , Junwei Yang , Runhui Huang , Jianhua Han , Lu Hou , Wei Zhang , Hang Xu

The swift progress of Multi-modal Large Models (MLLMs) has showcased their impressive ability to tackle tasks blending vision and language. Yet, most current models and benchmarks cater to scenarios with a narrow scope of visual and textual…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Chenyu Zhou , Mengdan Zhang , Peixian Chen , Chaoyou Fu , Yunhang Shen , Xiawu Zheng , Xing Sun , Rongrong Ji

Large language models (LLMs), such as ChatGPT, have demonstrated impressive capabilities in various tasks and attracted an increasing interest as a natural language interface across many domains. Recently, large vision-language models…

Computer Vision and Pattern Recognition · Computer Science 2025-09-01 Zhihao Chen , Bin Hu , Chuang Niu , Tao Chen , Yuxin Li , Hongming Shan , Ge Wang

Learning-based image quality assessment (IQA) has made remarkable progress in the past decade, but nearly all consider the two key components -- model and data -- in isolation. Specifically, model-centric IQA focuses on developing…

Computer Vision and Pattern Recognition · Computer Science 2023-12-11 Peibei Cao , Dingquan Li , Kede Ma

Blind image quality assessment (BIQA) remains challenging due to the diversity of distortion and image content variation, which complicate the distortion patterns crossing different scales and aggravate the difficulty of the regression…

Image and Video Processing · Electrical Eng. & Systems 2023-11-06 Qingyi Pan , Ning Guo , Letu Qingge , Jingyi Zhang , Pei Yang

Image Quality Assessment (IQA) models are increasingly deployed as perceptual critics to guide generative models and image restoration. This role demands not only accurate scores but also actionable, localized feedback. However, current…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Xudong Li , Jiaxi Tan , Ziyin Zhou , Yan Zhong , Zihao Huang , Jingyuan Zheng , Yan Zhang , Xiawu Zheng , Rongrong Ji

Despite impressive performance on standard benchmarks, multimodal large language models (MLLMs) face critical challenges in real-world clinical environments where medical images inevitably suffer various quality degradations. Existing…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Jiyao Liu , Junzhi Ning , Chenglong Ma , Wanying Qu , Jianghan Shen , Siqi Luo , Jinjie Wei , Jin Ye , Pengze Li , Tianbin Li , Jiashi Lin , Hongming Shan , Xinzhe Luo , Xiaohong Liu , Lihao Liu , Junjun He , Ningsheng Xu

Pairwise image quality assessment (IQA) in professional photography requires a model not only to identify the preferred image between two candidates, but also to provide convincing and image-grounded reasoning. In the NTIRE 2026 RAIM…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Xinli Yue , JianHui Sun , Tao Shao , Liangchao Yao , Fan Xia , Yuetang Deng

Multimodal large language models (MLLMs) have made significant strides by integrating visual and textual modalities. A critical factor in training MLLMs is the quality of image-text pairs within multimodal pretraining datasets. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Han Huang , Yuqi Huo , Zijia Zhao , Haoyu Lu , Shu Wu , Bingning Wang , Qiang Liu , Weipeng Chen , Liang Wang

Image Quality Assessment (IQA) is a core task in computer vision. Multimodal methods based on vision-language models, such as CLIP, have demonstrated exceptional generalization capabilities in IQA tasks. To address the issues of excessive…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Yongkang Hou , Jiarun Song

In the real world, knowledge often exists in a multimodal and heterogeneous form. Addressing the task of question answering with hybrid data types, including text, tables, and images, is a challenging task (MMHQA). Recently, with the rise…

Computation and Language · Computer Science 2023-09-12 Weihao Liu , Fangyu Lei , Tongxu Luo , Jiahe Lei , Shizhu He , Jun Zhao , Kang Liu

In-context learning (ICL) allows large models to adapt to tasks using a few examples, yet its extension to vision-language models (VLMs) remains fragile. Our analysis reveals that the fundamental limitation lies in an inductive gap, models…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Haoyu Wang , Haonan Wang , Yuyan Chen , Jun Chen , Gang Liu , Qian Wang , Jiahong Yan , Yanghua Xiao

Evaluating the performance of Multi-modal Large Language Models (MLLMs), integrating both point cloud and language, presents significant challenges. The lack of a comprehensive assessment hampers determining whether these models truly…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 Junjie Zhang , Tianci Hu , Xiaoshui Huang , Yongshun Gong , Dan Zeng

Recent advances of large multi-modality models (LMM) have greatly improved the ability of image quality assessment (IQA) method to evaluate and explain the quality of visual content. However, these advancements are mostly focused on overall…

Computer Vision and Pattern Recognition · Computer Science 2024-07-25 Chaofeng Chen , Sensen Yang , Haoning Wu , Liang Liao , Zicheng Zhang , Annan Wang , Wenxiu Sun , Qiong Yan , Weisi Lin

Document quality assessment is critical for a wide range of applications including document digitization, OCR, and archival. However, existing approaches often struggle to provide accurate and robust quality scores, limiting their…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Junjie Gao , Runze Liu , Yingzhe Peng , Shujian Yang , Jin Zhang , Kai Yang , Zhiyuan You

Visual Question Answering (VQA) with multiple choice questions enables a vision-centric evaluation of Multimodal Large Language Models (MLLMs). Although it reliably checks the existence of specific visual abilities, it is easier for the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Manu Gaur , Darshan Singh S , Makarand Tapaswi

Recent advances in image editing have heightened the need for reliable Image Editing Quality Assessment (IEQA). Unlike traditional methods, IEQA requires complex reasoning over multimodal inputs and multi-dimensional assessments. Existing…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Xinjie Zhang , Qiang Li , Xiaowen Ma , Axi Niu , Li Yan , Qingsen Yan