English
Related papers

Related papers: ServImage: An Image Generation and Editing Benchma…

200 papers

Portrait composition plays a central role in portrait aesthetics and visual communication, yet existing datasets and benchmarks mainly focus on coarse aesthetic scoring, generic image aesthetics, or unconstrained portrait generation. This…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Yuyang Sha , Zijie Lou , Youyun Tang , Xiaochao Qu , Zheng Qu , Ben Xia , Haoxiang Li , Ting Liu , Luoqi Liu

We address the task of evaluating image description generation systems. We propose a novel image-aware metric for this task: VIFIDEL. It estimates the faithfulness of a generated caption with respect to the content of the actual image,…

Computation and Language · Computer Science 2019-07-23 Pranava Madhyastha , Josiah Wang , Lucia Specia

Despite the remarkable progress of Vision-Language Models (VLMs) in adopting "Thinking-with-Images" capabilities, accurately evaluating the authenticity of their reasoning process remains a critical challenge. Existing benchmarks mainly…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Xuchen Li , Xuzhao Li , Renjie Pi , Shiyu Hu , Jian Zhao , Jiahui Gao

Large-scale product recognition is one of the major applications of computer vision and machine learning in the e-commerce domain. Since the number of products is typically much larger than the number of categories of products, image-based…

Computer Vision and Pattern Recognition · Computer Science 2021-07-14 Jiangbo Yuan , An-Ti Chiang , Wen Tang , Antonio Haro

Text-to-image generation using diffusion models has gained increasing popularity due to their ability to produce high-quality, realistic images based on text prompts. However, efficiently serving these models is challenging due to their…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-05-13 Sohaib Ahmad , Qizheng Yang , Haoliang Wang , Ramesh K. Sitaraman , Hui Guan

The evolution of video generation toward complex, multi-shot narratives has exposed a critical deficit in current evaluation methods. Existing benchmarks remain anchored to single-shot paradigms, lacking the comprehensive story assets and…

Multimedia · Computer Science 2026-03-02 Haoyuan Shi , Yunxin Li , Nanhao Deng , Zhenran Xu , Xinyu Chen , Longyue Wang , Baotian Hu , Min Zhang

Story visualization is an under-explored task that falls at the intersection of many important research directions in both computer vision and natural language processing. In this task, given a series of natural language captions which…

Computation and Language · Computer Science 2021-05-24 Adyasha Maharana , Darryl Hannan , Mohit Bansal

Multimodal large language models (LLMs) have demonstrated impressive capabilities in generating high-quality images from textual instructions. However, their performance in generating scientific images--a critical application for…

Sketching is a powerful artistic technique for capturing essential visual information about real-world objects and has increasingly attracted attention in image synthesis research. However, the field lacks a unified benchmark to evaluate…

Computer Vision and Pattern Recognition · Computer Science 2025-04-10 Xingyue Lin , Xingjian Hu , Shuai Peng , Jianhua Zhu , Liangcai Gao

How should we evaluate the quality of generative models? Many existing metrics focus on a model's producibility, i.e. the quality and breadth of outputs it can generate. However, the actual value from using a generative model stems not just…

Machine Learning · Computer Science 2025-11-13 Keyon Vafa , Sarah Bentley , Jon Kleinberg , Sendhil Mullainathan

Personalized dual-person portrait customization has considerable potential applications, such as preserving emotional memories and facilitating wedding photography planning. However, the absence of a benchmark dataset hinders the pursuit of…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Ting Pan , Ye Wang , Peiguang Jing , Rui Ma , Zili Yi , Yu Liu

Human image editing includes tasks like changing a person's pose, their clothing, or editing the image according to a text prompt. However, prior work often tackles these tasks separately, overlooking the benefit of mutual reinforcement…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Nannan Li , Qing Liu , Krishna Kumar Singh , Yilin Wang , Jianming Zhang , Bryan A. Plummer , Zhe Lin

This study assesses the ability of Large Vision-Language Models (LVLMs) to differentiate between AI-generated and human-generated images. It introduces a new automated benchmark construction method for this evaluation. The experiment…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Haokun Zhou , Yipeng Hong

Recent advances in multi-modal generative models have enabled significant progress in instruction-based image editing. However, while these models produce visually plausible outputs, their capacity for knowledge-based reasoning editing…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Yongliang Wu , Zonghui Li , Xinting Hu , Xinyu Ye , Xianfang Zeng , Gang Yu , Wenbo Zhu , Bernt Schiele , Ming-Hsuan Yang , Xu Yang

Machine-learning excels in many areas with well-defined goals. However, a clear goal is usually not available in art forms, such as photography. The success of a photograph is measured by its aesthetic value, a very subjective concept. This…

Computer Vision and Pattern Recognition · Computer Science 2017-07-13 Hui Fang , Meng Zhang

The vision and language generative models have been overgrown in recent years. For video generation, various open-sourced models and public-available services have been developed to generate high-quality videos. However, these methods often…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Yaofang Liu , Xiaodong Cun , Xuebo Liu , Xintao Wang , Yong Zhang , Haoxin Chen , Yang Liu , Tieyong Zeng , Raymond Chan , Ying Shan

Visual aesthetic assessment has been an active research field for decades. Although latest methods have achieved promising performance on benchmark datasets, they typically rely on a large number of manual annotations including both…

Computer Vision and Pattern Recognition · Computer Science 2019-12-04 Kekai Sheng , Weiming Dong , Menglei Chai , Guohui Wang , Peng Zhou , Feiyue Huang , Bao-Gang Hu , Rongrong Ji , Chongyang Ma

Unconditional human image generation is an important task in vision and graphics, which enables various applications in the creative industry. Existing studies in this field mainly focus on "network engineering" such as designing new…

Computer Vision and Pattern Recognition · Computer Science 2022-04-26 Jianglin Fu , Shikai Li , Yuming Jiang , Kwan-Yee Lin , Chen Qian , Chen Change Loy , Wayne Wu , Ziwei Liu

Lifestyle images are photographs that capture environments and objects in everyday settings. In furniture product marketing, advertisers often create lifestyle images containing products to resonate with potential buyers, allowing buyers to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Jialu Gao , Mithun Das Gupta , Qun Li , Raveena Kshatriya , Andrew D. Wilson , Keng-hao Chang , Balasaravanan Thoravi Kumaravel

Recent advances in multimodal large language models (MLLMs) have led to impressive progress across various benchmarks. However, their capability in understanding infrared images remains unexplored. To address this gap, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Tao Zhang , Yuyang Hong , Yang Xia , Kun Ding , Zeyu Zhang , Ying Wang , Shiming Xiang , Chunhong Pan
‹ Prev 1 4 5 6 7 8 10 Next ›