English
Related papers

Related papers: MMIG-Bench: Towards Comprehensive and Explainable …

200 papers

With the rapid advancement of video understanding, existing benchmarks are becoming increasingly saturated, exposing a critical discrepancy between inflated leaderboard scores and real-world model capabilities. To address this widening gap,…

Although recent large multimodal models (LMMs) demonstrate impressive progress on vision language tasks, their alignment with human centered (HC) principles, such as fairness, ethics, inclusivity, empathy, and robustness; remains poorly…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Shaina Raza , Aravind Narayanan , Vahid Reza Khazaie , Ashmal Vayani , Ahmed Y. Radwan , Mukund S. Chettiar , Amandeep Singh , Mubarak Shah , Deval Pandya

The progress in the generation of synthetic images has made it crucial to assess their quality. While several metrics have been proposed to assess the rendering of images, it is crucial for Text-to-Image (T2I) models, which generate images…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Paul Grimal , Hervé Le Borgne , Olivier Ferret , Julien Tourille

Multimodal large language models (MLLMs), building upon the foundation of powerful large language models (LLMs), have recently demonstrated exceptional capabilities in generating not only texts but also images given interleaved multimodal…

Computer Vision and Pattern Recognition · Computer Science 2023-11-30 Bohao Li , Yuying Ge , Yixiao Ge , Guangzhi Wang , Rui Wang , Ruimao Zhang , Ying Shan

Recent advances in text-to-image (T2I) generation have achieved remarkable visual outcomes through large-scale rectified flow models. However, how these models behave under long prompts remains underexplored. Long prompts encode rich…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Bo-Kai Ruan , Teng-Fang Hsiao , Ling Lo , Yi-Lun Wu , Hong-Han Shuai

Generating high-quality images without prompt engineering expertise remains a challenge for text-to-image (T2I) models, which often misinterpret poorly structured prompts, leading to distortions and misalignments. While humans easily…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Nisan Chhetri , Arpan Sainju

Recent advances in image generation, often driven by proprietary systems like GPT-4o Image Gen, regularly introduce new capabilities that reshape how users interact with these models. Existing benchmarks often lag behind and fail to capture…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Jiaxin Ge , Grace Luo , Heekyung Lee , Nishant Malpani , Long Lian , XuDong Wang , Aleksander Holynski , Trevor Darrell , Sewon Min , David M. Chan

Generative diffusion models are developing rapidly and attracting increasing attention due to their wide range of applications. Image-to-Video (I2V) generation has become a major focus in the field of video synthesis. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Ailing Zhang , Lina Lei , Dehong Kong , Zhixin Wang , Jiaqi Xu , Fenglong Song , Chun-Le Guo , Chang Liu , Fan Li , Jie Chen

Recent advances in multi-modal generative models have driven substantial improvements in image editing. However, current generative models still struggle with handling diverse and complex image editing tasks that require implicit reasoning,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Feng Han , Yibin Wang , Chenglin Li , Zheming Liang , Dianyi Wang , Yang Jiao , Zhipeng Wei , Chao Gong , Cheng Jin , Jingjing Chen , Jiaqi Wang

Unified multimodal models target joint understanding, reasoning, and generation, but current image editing benchmarks are largely confined to natural images and shallow commonsense reasoning, offering limited assessment of this capability…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Mingxin Liu , Ziqian Fan , Zhaokai Wang , Leyao Gu , Zirun Zhu , Yiguo He , Yuchen Yang , Changyao Tian , Xiangyu Zhao , Ning Liao , Shaofeng Zhang , Qibing Ren , Zhihang Zhong , Xuanhe Zhou , Junchi Yan , Xue Yang

Text-to-image generation models have achieved strong performance in culturally homogeneous settings, yet their ability to generate multicultural scenes, where people and landmarks originate from different cultures, remains largely…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Parth Bhalerao , Mounika Yalamarty , Brian Trinh , Oana Ignat

Infographics are composite visual artifacts that combine data visualizations with textual and illustrative elements to communicate information. While recent text-to-image (T2I) models can generate aesthetically appealing images, their…

Video generation has witnessed significant advancements, yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Ziqi Huang , Yinan He , Jiashuo Yu , Fan Zhang , Chenyang Si , Yuming Jiang , Yuanhan Zhang , Tianxing Wu , Qingyang Jin , Nattapol Chanpaisit , Yaohui Wang , Xinyuan Chen , Limin Wang , Dahua Lin , Yu Qiao , Ziwei Liu

Reward models play an essential role in training vision-language models (VLMs) by assessing output quality to enable aligning with human preferences. Despite their importance, the research community lacks comprehensive open benchmarks for…

Computer Vision and Pattern Recognition · Computer Science 2025-02-21 Michihiro Yasunaga , Luke Zettlemoyer , Marjan Ghazvininejad

In the quest for artificial general intelligence, Multi-modal Large Language Models (MLLMs) have emerged as a focal point in recent advancements. However, the predominant focus remains on developing their capabilities in static image…

In this paper, we present an empirical study introducing a nuanced evaluation framework for text-to-image (T2I) generative models, applied to human image synthesis. Our framework categorizes evaluations into two distinct groups: first,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Muxi Chen , Yi Liu , Jian Yi , Changran Xu , Qiuxia Lai , Hongliang Wang , Tsung-Yi Ho , Qiang Xu

The rapid advancement of multimodal large language models (MLLMs) has profoundly impacted the document domain, creating a wide array of application scenarios. This progress highlights the need for a comprehensive benchmark to evaluate these…

Computation and Language · Computer Science 2025-05-23 Siqi Li , Yufan Shen , Xiangnan Chen , Jiayi Chen , Hengwei Ju , Haodong Duan , Song Mao , Hongbin Zhou , Bo Zhang , Bin Fu , Pinlong Cai , Licheng Wen , Botian Shi , Yong Liu , Xinyu Cai , Yu Qiao

In this work, we study the problem of generating novel images from complex multimodal prompt sequences. While existing methods achieve promising results for text-to-image generation, they often struggle to capture fine-grained details from…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Amandeep Kumar , Muzammal Naseer , Sanath Narayan , Rao Muhammad Anwer , Salman Khan , Hisham Cholakkal

The remarkable progress of Multi-modal Large Language Models (MLLMs) has attracted significant attention due to their superior performance in visual contexts. However, their capabilities in turning visual figure to executable code, have not…

Computation and Language · Computer Science 2024-05-14 Chengyue Wu , Yixiao Ge , Qiushan Guo , Jiahao Wang , Zhixuan Liang , Zeyu Lu , Ying Shan , Ping Luo

We introduce ``Idea to Image,'' a system that enables multimodal iterative self-refinement with GPT-4V(ision) for automatic image design and generation. Humans can quickly identify the characteristics of different text-to-image (T2I) models…

Computer Vision and Pattern Recognition · Computer Science 2024-08-15 Zhengyuan Yang , Jianfeng Wang , Linjie Li , Kevin Lin , Chung-Ching Lin , Zicheng Liu , Lijuan Wang