中文
相关论文

相关论文: M3-AGIQA: Multimodal, Multi-Round, Multi-Aspect AI…

200 篇论文

Omnidirectional image quality assessment (OIQA) has been widely investigated in the past few years and achieved much success. However, most of existing studies are dedicated to solve the uniform distortion problem in OIQA, which has a…

图像与视频处理 · 电气工程与系统科学 2025-01-22 Jiebin Yan , Jiale Rao , Junjie Chen , Ziwen Tan , Weide Liu , Yuming Fang

The growing capabilities of AI in generating video content have brought forward significant challenges in effectively evaluating these videos. Unlike static images or text, video content involves complex spatial and temporal dynamics which…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Xiao Liu , Xinhao Xiang , Zizhong Li , Yongheng Wang , Zhuoheng Li , Zhuosheng Liu , Weidi Zhang , Weiqi Ye , Jiawei Zhang

In no-reference image quality assessment (NR-IQA), the challenge of limited dataset sizes hampers the development of robust and generalizable models. Conventional methods address this issue by utilizing large datasets to extract rich…

计算机视觉与模式识别 · 计算机科学 2024-10-08 Daekyu Kwon , Dongyoung Kim , Sehwan Ki , Younghyun Jo , Hyong-Euk Lee , Seon Joo Kim

Image Quality Assessment (IQA) and Image Aesthetic Assessment (IAA) aim to simulate human subjective perception of image visual quality and aesthetic appeal. Despite distinct learning objectives, they have underlying interconnectedness due…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Hantao Zhou , Longxiang Tang , Rui Yang , Guanyi Qin , Yan Zhang , Yutao Li , Xiu Li , Runze Hu , Guangtao Zhai

Image quality assessment (IQA) represents a pivotal challenge in image-focused technologies, significantly influencing the advancement trajectory of image processing and computer vision. Recently, IQA has witnessed a notable surge in…

计算机视觉与模式识别 · 计算机科学 2025-02-13 Chengqian Ma , Zhengyi Shi , Zhiqiang Lu , Shenghao Xie , Fei Chao , Yao Sui

We introduce AEGIS, A holistic benchmark for Evaluating forensic analysis of AI-Generated academic ImageS. Compared to existing benchmarks, AEGIS features three key advances: (1) Domain-Specific Complexity: covering seven academic…

AI-based text-to-image models do not only excel at generating realistic images, they also give designers more and more fine-grained control over the image content. Consequently, these approaches have gathered increased attention within the…

计算机视觉与模式识别 · 计算机科学 2025-01-30 Sebastian Hartwig , Dominik Engel , Leon Sick , Hannah Kniesel , Tristan Payer , Poonam Poonam , Michael Glöckler , Alex Bäuerle , Timo Ropinski

In the early stages of architectural design, shoebox models are typically used as a simplified representation of building structures but require extensive operations to transform them into detailed designs. Generative artificial…

图形学 · 计算机科学 2025-03-06 Xusheng Du , Ruihan Gui , Zhengyang Wang , Ye Zhang , Haoran Xie

Blind image quality assessment (BIQA) is a task that predicts the perceptual quality of an image without its reference. Research on BIQA attracts growing attention due to the increasing amount of user-generated images and emerging mobile…

图像与视频处理 · 电气工程与系统科学 2023-03-24 Zhanxuan Mei , Yun-Cheng Wang , Xingze He , Yong Yan , C. -C. Jay Kuo

As recent multi-modality large language models (MLLMs) have shown formidable proficiency on various complex tasks, there has been increasing attention on debating whether these models could eventually mirror human intelligence. However,…

人工智能 · 计算机科学 2024-06-17 Wei Song , Yadong Li , Jianhua Xu , Guowei Wu , Lingfeng Ming , Kexin Yi , Weihua Luo , Houyi Li , Yi Du , Fangda Guo , Kaicheng Yu

With the increasing demand for image-based applications, the efficient and reliable evaluation of image quality has increased in importance. Measuring the image quality is of fundamental importance for numerous image processing…

多媒体 · 计算机科学 2014-07-01 Pedram Mohammadi , Abbas Ebrahimi-Moghadam , Shahram Shirani

We present M$^3$-VQA, a novel knowledge-based Visual Question Answering (VQA) benchmark, to enhance the evaluation of multimodal large language models (MLLMs) in fine-grained multimodal entity understanding and complex multi-hop reasoning.…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Jiatong Ma , Longteng Guo , Yuchen Liu , Zijia Zhao , Dongze Hao , Xuanxu Lin , Jing Liu

Multimodal Large Language Models (MLLMs) have shown impressive abilities in understanding and reasoning over conventional images. However, their perception of 360{\deg} images remains largely underexplored. Unlike conventional images,…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Huyen T. T. Tran , Van-Quang Nguyen , Farros Alferro , Kang-Jun Liu , Takayuki Okatani

The rapid proliferation of AI-Generated Images (AIGIs) has introduced severe risks of misinformation, making AIGI detection a critical yet challenging task. While traditional detection paradigms mainly rely on low-level features, recent…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Chenyang Zhu , Maorong Wang , Jun Liu , Ching-Chun Chang , Isao Echizen

Image quality assessment (IQA) serves as the golden standard for all models' performance in nearly all computer vision fields. However, it still suffers from poor out-of-distribution generalization ability and expensive training costs. To…

计算机视觉与模式识别 · 计算机科学 2024-10-11 Kai Liu , Ziqing Zhang , Wenbo Li , Renjing Pei , Fenglong Song , Xiaohong Liu , Linghe Kong , Yulun Zhang

The rapid progress of generative artificial intelligence has exposed fundamental limitations in existing evaluation methodologies, particularly for open-ended, creative, and human-facing tasks. Traditional automatic metrics rely on…

人工智能 · 计算机科学 2026-05-19 Marjan Veysi , Pirooz Shamsinejadbabaki , Mohammad Zare , Mohammad Sabouri

Visual-textual inconsistency (VTI) evaluation plays a crucial role in cleansing vision-language data. Its main challenges stem from the high variety of image captioning datasets, where differences in content can create a range of…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Zihao Zhu , Hongbao Zhang , Guanzong Wu , Siwei Lyu , Baoyuan Wu

Recent advances in Image Quality Assessment (IQA) have leveraged Multi-modal Large Language Models (MLLMs) to generate descriptive explanations. However, despite their strong visual perception modules, these models often fail to reliably…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Yuan Li , Zitang Sun , Yen-Ju Chen , Shin'ya Nishida

No-Reference Image Quality Assessment (NR-IQA) aims to assess the perceptual quality of images in accordance with human subjective perception. Unfortunately, existing NR-IQA methods are far from meeting the needs of predicting accurate…

计算机视觉与模式识别 · 计算机科学 2022-04-22 Sidi Yang , Tianhe Wu , Shuwei Shi , Shanshan Lao , Yuan Gong , Mingdeng Cao , Jiahao Wang , Yujiu Yang

In this paper, we propose XGC-AVis, a multi-agent framework that enhances the audio-video temporal alignment capabilities of multimodal large models (MLLMs) and improves the efficiency of retrieving key video segments through 4 stages:…

多媒体 · 计算机科学 2025-09-30 Yuqin Cao , Xiongkuo Min , Yixuan Gao , Wei Sun , Zicheng Zhang , Jinliang Han , Guangtao Zhai