中文
相关论文

相关论文: Q-DeepSight: Incentivizing Thinking with Images fo…

200 篇论文

Large language models perform well on many medical QA benchmarks, but real clinical reasoning often requires integrating evidence across multiple images rather than interpreting a single view. We introduce MedThinkVQA, an expert-annotated…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Zonghai Yao , Benlu Wang , Yifan Zhang , Junda Wang , Iris Xia , Zhipeng Tang , Shuo Han , Feiyun Ouyang , Zhichao Yang , Arman Cohan , Hong Yu

Current full-reference image quality assessment (FR-IQA) methods often fuse features from reference and distorted images, overlooking that color and luminance distortions occur mainly at low frequencies, whereas edge and texture distortions…

图像与视频处理 · 电气工程与系统科学 2024-12-23 Xuekai Wei , Junyu Zhang , Qinlin Hu , Mingliang Zhou\\Yong Feng , Weizhi Xian , Huayan Pu , Sam Kwong

Image Quality Assessment algorithms predict a quality score for a pristine or distorted input image, such that it correlates with human opinion. Traditional methods required a non-distorted "reference" version of the input image to compare…

图像与视频处理 · 电气工程与系统科学 2020-07-21 Subhayan Mukherjee , Giuseppe Valenzise , Irene Cheng

Image quality assessment (IQA) algorithms aim to reproduce the human's perception of the image quality. The growing popularity of image enhancement, generation, and recovery models instigated the development of many methods to assess their…

图像与视频处理 · 电气工程与系统科学 2023-02-17 Segrey Kastryulin , Jamil Zakirov , Nicola Pezzotti , Dmitry V. Dylov

Deep learning based image quality assessment (IQA) models usually learn to predict image quality from a single dataset, leading the model to overfit specific scenes. To account for this, mixed datasets training can be an effective way to…

计算机视觉与模式识别 · 计算机科学 2022-11-15 Zhaopeng Feng , Keyang Zhang , Shuyue Jia , Baoliang Chen , Shiqi Wang

Image quality assessment (IQA) serves as the golden standard for all models' performance in nearly all computer vision fields. However, it still suffers from poor out-of-distribution generalization ability and expensive training costs. To…

计算机视觉与模式识别 · 计算机科学 2024-10-11 Kai Liu , Ziqing Zhang , Wenbo Li , Renjing Pei , Fenglong Song , Xiaohong Liu , Linghe Kong , Yulun Zhang

We present Thinking with Generated Images, a novel paradigm that fundamentally transforms how large multimodal models (LMMs) engage with visual reasoning by enabling them to natively think across text and vision modalities through…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Ethan Chern , Zhulin Hu , Steffi Chern , Siqi Kou , Jiadi Su , Yan Ma , Zhijie Deng , Pengfei Liu

Perceptual image restoration seeks for high-fidelity images that most likely degrade to given images. For better visual quality, previous work proposed to search for solutions within the natural image manifold, by exploiting the latent…

图像与视频处理 · 电气工程与系统科学 2021-03-05 Chaoyi Han , Yiping Duan , Xiaoming Tao , Jianhua Lu

Image Quality Assessment (IQA) is a core task in computer vision. Multimodal methods based on vision-language models, such as CLIP, have demonstrated exceptional generalization capabilities in IQA tasks. To address the issues of excessive…

计算机视觉与模式识别 · 计算机科学 2025-07-24 Yongkang Hou , Jiarun Song

With the rapid evolution of the Text-to-Image (T2I) model in recent years, their unsatisfactory generation result has become a challenge. However, uniformly refining AI-Generated Images (AIGIs) of different qualities not only limited…

计算机视觉与模式识别 · 计算机科学 2024-01-03 Chunyi Li , Haoning Wu , Zicheng Zhang , Hongkun Hao , Kaiwei Zhang , Lei Bai , Xiaohong Liu , Xiongkuo Min , Weisi Lin , Guangtao Zhai

Image Quality Assessment (IQA) has progressed from scalar quality prediction to more interpretable, human-aligned evaluation paradigms. In this work, we address the emerging challenge of detailed and explainable IQA by proposing iDETEX-a…

计算机视觉与模式识别 · 计算机科学 2025-10-21 Zhaoran Zhao , Xinli Yue , Jianhui Sun , Yuhao Xie , Tao Shao , Liangchao Yao , Fan Xia , Yuetang Deng

Reasoning-based image quality assessment (IQA) models trained through reinforcement learning (RL) exhibit exceptional generalization, yet the underlying mechanisms and critical factors driving this capability remain underexplored in current…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Shijie Zhao , Xuanyu Zhang , Weiqi Li , Junlin Li , Li Zhang , Tianfan Xue , Jian Zhang

Recent advances in multimodal language models (MLLMs) have made thinking with images a dominant paradigm for multimodal reasoning. However, existing methods still fail to ensure evidence-answer consistency, where correct answers must be…

人工智能 · 计算机科学 2026-05-22 Tianrun Xu , Haoda Jing , Ye Li , Yuquan Wei , Jun Feng , Guanyu Chen , Haichuan Gao , Tianren Zhang , Feng Chen

We present IQA-Spider, the first image quality assessment (IQA) framework that unifies reasoning, grounding, and referring into a single LMM-based framework for multi-granularity quality understanding. Existing LMM-based IQA methods…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Xinge Peng , Yiting Lu , Xin Li , Zhibo Chen

Multimodal large language models (MLLMs) have achieved impressive performance across various tasks such as image captioning and visual question answer(VQA); however, they often struggle to accurately interpret depth information inherent in…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Hao Yang , Hongbo Zhang , Yanyan Zhao , Bing Qin

Existing full-reference image quality assessment (FR-IQA) methods often fail to capture the complex causal mechanisms that underlie human perceptual responses to image distortions, limiting their ability to generalize across diverse…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Wenhao Shen , Mingliang Zhou , Yu Chen , Xuekai Wei , Jun Luo , Huayan Pu , Weijia Jia

Traditional Image Quality Assessment (IQA) metrics typically fall into one of two extremes: rigid, hand-crafted mathematical models or "black-box" deep learning architectures that completely lack interpretability. To bridge this gap, we…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Ruchika Gupta , Illya Bakurov , Nathan Haut , Wolfgang Banzhaf

Existing deep network-based full-reference image quality assessment (FR-IQA) models typically work by performing pairwise comparisons of deep features from the reference and distorted images. In this paper, we approach this problem from a…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Zhen Zhang , Jielei Chu , Tian Zhang , Lin Ma , Fengmao Lv , Weide Liu , Tianrui Li , Yuming Fang

Automatic perception of image quality is a challenging problem that impacts billions of Internet and social media users daily. To advance research in this field, we propose a no-reference image quality assessment (NR-IQA) method termed…

计算机视觉与模式识别 · 计算机科学 2024-05-08 Zhen Zhang

Instruction-based image editing has emerged as a prominent research area, which, benefiting from image generation foundation models, have achieved high aesthetic quality, making instruction-following capability the primary challenge.…

计算机视觉与模式识别 · 计算机科学 2025-12-08 Hongyu Li , Manyuan Zhang , Dian Zheng , Ziyu Guo , Yimeng Jia , Kaituo Feng , Hao Yu , Yexin Liu , Yan Feng , Peng Pei , Xunliang Cai , Linjiang Huang , Hongsheng Li , Si Liu