English
Related papers

Related papers: HEIE: MLLM-Based Hierarchical Explainable AIGC Ima…

200 papers

Perceiving visual semantics embedded within consecutive characters is a crucial yet under-explored capability for both Large Language Models (LLMs) and Multi-modal Large Language Models (MLLMs). In this work, we select ASCII art as a…

Computation and Language · Computer Science 2025-09-26 Qi Jia , Xiang Yue , Shanshan Huang , Ziheng Qin , Yizhu Liu , Bill Yuchen Lin , Yang You , Guangtao Zhai

Multi-modal large language models (MLLMs), such as GPT-4o, excel at integrating text and visual data but face systematic challenges when interpreting ambiguous or incomplete visual stimuli. This study leverages statistical modeling to…

Machine Learning · Computer Science 2024-12-09 Ching-Yi Wang

Knowledge Tracing (KT) aims to mine students' evolving knowledge states and predict their future question-answering performance. Existing methods based on heterogeneous information networks (HINs) are prone to introducing noises due to…

Artificial Intelligence · Computer Science 2025-11-20 Zhiyi Duan , Zixing Shi , Hongyu Yuan , Qi Wang

High-resolution and variable-shape images have not yet been properly addressed by the AI community. The approach of down-sampling data often used with convolutional neural networks is sub-optimal for many tasks, and has too many drawbacks…

Computer Vision and Pattern Recognition · Computer Science 2020-09-30 Ferran Parés , Dario Garcia-Gasulla , Harald Servat , Jesús Labarta , Eduard Ayguadé

Recent large vision-language models (LVLMs) have advanced capabilities in visual question answering (VQA). However, interpreting where LVLMs direct their visual attention remains a significant challenge, yet is essential for understanding…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Guanxi Shen

Current instruction-based image editing (IBIE) methods struggle with challenging editing tasks, as both editing types and sample counts of existing datasets are limited. Moreover, traditional dataset construction often contains noisy…

Computer Vision and Pattern Recognition · Computer Science 2025-09-19 Mingsong Li , Lin Liu , Hongjun Wang , Haoxing Chen , Xijun Gu , Shizhan Liu , Dong Gong , Junbo Zhao , Zhenzhong Lan , Jianguo Li

With the advancement of technology for artificial intelligence (AI) based solutions and analytics compute engines, machine learning (ML) models are getting more complex day by day. Most of these models are generally used as a black box…

Machine Learning · Computer Science 2022-10-11 P. Sai Ram Aditya , Mayukha Pal

Deep learning has recently gained popularity in digital pathology due to its high prediction quality. However, the medical domain requires explanation and insight for a better understanding beyond standard quantitative performance…

Image and Video Processing · Electrical Eng. & Systems 2020-04-28 Miriam Hägele , Philipp Seegerer , Sebastian Lapuschkin , Michael Bockmayr , Wojciech Samek , Frederick Klauschen , Klaus-Robert Müller , Alexander Binder

Current AI-Generated Image (AIGI) detection approaches predominantly rely on binary classification to distinguish real from synthetic images, often lacking interpretable or convincing evidence to substantiate their decisions. This…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Yao Xiao , Weiyan Chen , Jiahao Chen , Zijie Cao , Weijian Deng , Binbin Yang , Ziyi Dong , Xiangyang Ji , Wei Ke , Pengxu Wei , Liang Lin

Recent advances in instruction-based image editing have shown remarkable progress. However, existing methods remain limited to relatively simple editing operations, hindering real-world applications that require complex and compositional…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Xuehai Bai , Xiaoling Gu , Akide Liu , Hangjie Yuan , YiFan Zhang , Jack Ma

Unpaired Medical Image Enhancement (UMIE) aims to transform a low-quality (LQ) medical image into a high-quality (HQ) one without relying on paired images for training. While most existing approaches are based on Pix2Pix/CycleGAN and are…

Image and Video Processing · Electrical Eng. & Systems 2023-07-18 Chunming He , Kai Li , Guoxia Xu , Jiangpeng Yan , Longxiang Tang , Yulun Zhang , Xiu Li , Yaowei Wang

The field of explainable AI (XAI) has quickly become a thriving and prolific community. However, a silent, recurrent and acknowledged issue in this area is the lack of consensus regarding its terminology. In particular, each new…

Artificial Intelligence · Computer Science 2021-11-03 Sebastian Palacio , Adriano Lucieri , Mohsin Munir , Jörn Hees , Sheraz Ahmed , Andreas Dengel

In-context learning (ICL) enables multimodal large language models (MLLMs) to classify images from a few labelled examples. Yet, how these models use the provided context remains opaque. While Chain-of-Thought prompting is widely used,…

Artificial Intelligence · Computer Science 2026-05-28 Carmen Quiles-Ramírez , Leticia L. Rodríguez , Nicolás Martorell , Natalia Díaz-Rodríguez

While Large Language Models (LLMs) have demonstrated significant advancements in reasoning and agent-based problem-solving, current evaluation methodologies fail to adequately assess their capabilities: existing benchmarks either rely on…

The increasing complexity of LLMs presents significant challenges to their transparency and interpretability, necessitating the use of eXplainable AI (XAI) techniques to enhance trustworthiness and usability. This study introduces a…

Computation and Language · Computer Science 2025-04-09 Melkamu Abay Mersha , Mesay Gemeda Yigezu , Hassan Shakil , Ali K. AlShami , Sanghyun Byun , Jugal Kalita

The significant advancements in visual understanding and instruction following from Multimodal Large Language Models (MLLMs) have opened up more possibilities for broader applications in diverse and universal human-centric scenarios.…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Keliang Li , Zaifei Yang , Jiahe Zhao , Hongze Shen , Ruibing Hou , Hong Chang , Shiguang Shan , Xilin Chen

Nowadays, multimedia forensics faces unprecedented challenges due to the rapid advancement of multimedia generation technology thereby making Image Manipulation Localization (IML) crucial in the pursuit of truth. The key to IML lies in…

Computer Vision and Pattern Recognition · Computer Science 2023-10-11 Xiaochen Ma , Jizhe Zhou , Xiong Xu , Zhuohang Jiang , Chi-Man Pun

While Multimodal Large Language Models (MLLMs) have experienced significant advancement in visual understanding and reasoning, their potential to serve as powerful, flexible, interpretable, and text-driven models for Image Quality…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Tianhe Wu , Kede Ma , Jie Liang , Yujiu Yang , Lei Zhang

Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities. However, evaluating their capacity for human-like understanding in One-Image Guides remains insufficiently explored. One-Image Guides are…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Jiancong Xie , Wenjin Wang , Zhuomeng Zhang , Zihan Liu , Qi Liu , Ke Feng , Zixun Sun , Yuedong Yang

With the growing adoption of AI-based systems across everyday life, the need to understand their decision-making mechanisms is correspondingly increasing. The level at which we can trust the statistical inferences made from AI-based…

Machine Learning · Statistics 2024-04-15 Adam Spannaus , Heidi A. Hanson , Lynne Penberthy , Georgia Tourassi