English
Related papers

Related papers: IE-Bench: Advancing the Measurement of Text-Driven…

200 papers

With the rapid advancement of Vision Language Models (VLMs), VLM-based Image Quality Assessment (IQA) seeks to describe image quality linguistically to align with human expression and capture the multifaceted nature of IQA tasks. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Zhiyuan You , Jinjin Gu , Xin Cai , Zheyuan Li , Kaiwen Zhu , Chao Dong , Tianfan Xue

With the rapid advancements in AI-Generated Content (AIGC), AI-Generated Images (AIGIs) have been widely applied in entertainment, education, and social media. However, due to the significant variance in quality among different AIGIs, there…

Computer Vision and Pattern Recognition · Computer Science 2024-04-05 Chunyi Li , Tengchuan Kou , Yixuan Gao , Yuqin Cao , Wei Sun , Zicheng Zhang , Yingjie Zhou , Zhichao Zhang , Weixia Zhang , Haoning Wu , Xiaohong Liu , Xiongkuo Min , Guangtao Zhai

The rapid development and reduced barriers to entry for Text-to-Image (T2I) models have raised concerns about the biases in their outputs, but existing research lacks a holistic definition and evaluation framework of biases, limiting the…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Hanjun Luo , Ziye Deng , Ruizhe Chen , Zuozhu Liu

Image captioning models generally lack the capability to take into account user interest, and usually default to global descriptions that try to balance readability, informativeness, and information overload. On the other hand, VQA models…

Computer Vision and Pattern Recognition · Computer Science 2021-11-12 Edwin G. Ng , Bo Pang , Piyush Sharma , Radu Soricut

We propose T2I-ReasonBench, a benchmark evaluating reasoning capabilities of text-to-image (T2I) models. It consists of four dimensions: Idiom Interpretation, Textual Image Design, Entity-Reasoning and Scientific-Reasoning. We propose a…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Kaiyue Sun , Rongyao Fang , Chengqi Duan , Xian Liu , Xihui Liu

Recent progress in large-scale pre-training has led to the development of advanced vision-language models (VLMs) with remarkable proficiency in comprehending and generating multimodal content. Despite the impressive ability to perform…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Hang Hua , Jing Shi , Kushal Kafle , Simon Jenni , Daoan Zhang , John Collomosse , Scott Cohen , Jiebo Luo

Generative adversarial networks (GANs) have achieved impressive results today, but not all generated images are perfect. A number of quantitative criteria have recently emerged for generative model, but none of them are designed for a…

Image and Video Processing · Electrical Eng. & Systems 2020-07-15 Shuyang Gu , Jianmin Bao , Dong Chen , Fang Wen

Recently, text-guided image manipulation has received increasing attention in the research field of multimedia processing and computer vision due to its high flexibility and controllability. Its goal is to semantically manipulate parts of…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Ryugo Morita , Zhiqiang Zhang , Man M. Ho , Jinjia Zhou

Text-to-image diffusion models can generate diverse, high-fidelity images based on user-provided text prompts. Recent research has extended these models to support text-guided image editing. While text guidance is an intuitive editing…

Computer Vision and Pattern Recognition · Computer Science 2023-05-26 Jooyoung Choi , Yunjey Choi , Yunji Kim , Junho Kim , Sungroh Yoon

The increasing availability of image-text pairs has largely fueled the rapid advancement in vision-language foundation models. However, the vast scale of these datasets inevitably introduces significant variability in data quality, which…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Lei Zhang , Fangxun Shu , Tianyang Liu , Sucheng Ren , Hao Jiang , Cihang Xie

AEC drawings encode geometry and semantics through symbols, layout conventions, and dense annotation, yet it remains unclear whether modern multimodal and vision-language models can reliably interpret this graphical language. We present…

Artificial Intelligence · Computer Science 2026-01-09 Aleksei Kondratenko , Mussie Birhane , Houssame E. Hsain , Guido Maciocci

Visual-prompt-guided edit transfer aims to learn image transformations directly from example pairs, offering more precise and controllable editing than purely text-driven approaches. However, existing diffusion transformer-based methods…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Lan Chen , Qi Mao , Yiren Song , Yuchao Gu , Siwei Ma

Automatic image aesthetics assessment is a computer vision problem dealing with categorizing images into different aesthetic levels. The categorization is usually done by analyzing an input image and computing some measure of the degree to…

Computer Vision and Pattern Recognition · Computer Science 2022-02-08 Abbas Anwar , Saira Kanwal , Muhammad Tahir , Muhammad Saqib , Muhammad Uzair , Mohammad Khalid Imam Rahmani , Habib Ullah

It is well-known that there is no universal metric for image quality evaluation. In this case, distortion-specific metrics can be more reliable. The artifact imposed by image compression can be considered as a combination of various…

Image and Video Processing · Electrical Eng. & Systems 2024-02-05 S. Farhad Hosseini-Benvidi , Hossein Motamednia , Azadeh Mansouri , Mohammadreza Raei , Ahmad Mahmoudi-Aznaveh

Diffusion models for Text-to-Image (T2I) conditional generation have recently achieved tremendous success. Yet, aligning these models with user's intentions still involves a laborious trial-and-error process, and this challenging alignment…

Machine Learning · Computer Science 2025-02-12 Chao Wang , Giulio Franzese , Alessandro Finamore , Massimo Gallo , Pietro Michiardi

Studies have been conducted to prevent specific concepts from being generated from pretrained text-to-image generative models, achieving concept erasure in various ways. However, the performance evaluation of these studies is still largely…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Masane Fuchi , Tomohiro Takagi

Image Quality Assessment (IQA) based on human subjective preferences has undergone extensive research in the past decades. However, with the development of communication protocols, the visual data consumption volume of machines has…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Chunyi Li , Yuan Tian , Xiaoyue Ling , Zicheng Zhang , Haodong Duan , Haoning Wu , Ziheng Jia , Xiaohong Liu , Xiongkuo Min , Guo Lu , Weisi Lin , Guangtao Zhai

Real-world design tasks - such as picture book creation, film storyboard development using character sets, photo retouching, visual effects, and font transfer - are highly diverse and complex, requiring deep interpretation and extraction of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Chen Liang , Lianghua Huang , Jingwu Fang , Huanzhang Dou , Wei Wang , Zhi-Fan Wu , Yupeng Shi , Junge Zhang , Xin Zhao , Yu Liu

The rapid advancement of Large Multi-modal Foundation Models (LMM) has paved the way for the possible Explainable Image Quality Assessment (EIQA) with instruction tuning from two perspectives: overall quality explanation, and attribute-wise…

Computer Vision and Pattern Recognition · Computer Science 2025-04-03 Yiting Lu , Xin Li , Haoning Wu , Bingchen Li , Weisi Lin , Zhibo Chen

Recently, AIGC image quality assessment (AIGCIQA), which aims to assess the quality of AI-generated images (AIGIs) from a human perception perspective, has emerged as a new topic in computer vision. Unlike common image quality assessment…

Computer Vision and Pattern Recognition · Computer Science 2024-01-12 Jiquan Yuan , Xinyan Cao , Jinming Che , Qinyuan Wang , Sen Liang , Wei Ren , Jinlong Lin , Xixin Cao