English
Related papers

Related papers: The ArtBench Dataset: Benchmarking Generative Mode…

200 papers

Multimodal Large Language Models (MLLM) have made significant progress in the field of document analysis. Despite this, existing benchmarks typically focus only on extracting text and simple layout information, neglecting the complex…

Computer Vision and Pattern Recognition · Computer Science 2024-07-04 Lei Chen , Feng Yan , Yujie Zhong , Shaoxiang Chen , Zequn Jie , Lin Ma

Synthesizing realistic images from human drawn sketches is a challenging problem in computer graphics and vision. Existing approaches either need exact edge maps, or rely on retrieval of existing photographs. In this work, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2018-04-16 Wengling Chen , James Hays

Recent advances in generative modeling can create remarkably realistic synthetic videos, making it increasingly difficult for humans to distinguish them from real ones and necessitating reliable detection methods. However, two key…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Long Ma , Zihao Xue , Yan Wang , Zhiyuan Yan , Jin Xu , Xiaorui Jiang , Haiyang Yu , Yong Liao , Zhen Bi

Multimodal learning has gained attention for its capacity to integrate information from different modalities. However, it is often hindered by the multimodal imbalance problem, where certain modality dominates while others remain…

Machine Learning · Computer Science 2025-06-16 Shaoxuan Xu , Menglu Cui , Chengxiang Huang , Hongfa Wang , Di Hu

Hand-drawn sketches are a natural and efficient medium for capturing and conveying ideas. Despite significant advancements in controllable natural image generation, translating freehand sketches into structured, machine-readable diagrams…

Artificial Intelligence · Computer Science 2025-08-05 Cheng Tan , Qi Chen , Jingxuan Wei , Gaowei Wu , Zhangyang Gao , Siyuan Li , Bihui Yu , Ruifeng Guo , Stan Z. Li

In recent years, advancements in generative artificial intelligence have led to the development of sophisticated tools capable of mimicking diverse artistic styles, opening new possibilities for digital creativity and artistic expression.…

Computer Vision and Pattern Recognition · Computer Science 2025-02-26 Andrea Asperti , Franky George , Tiberio Marras , Razvan Ciprian Stricescu , Fabio Zanotti

The artistic style of a painting is a rich descriptor that reveals both visual and deep intrinsic knowledge about how an artist uniquely portrays and expresses their creative vision. Accurate categorization of paintings across different…

Computer Vision and Pattern Recognition · Computer Science 2020-12-08 Akshay Joshi , Ankit Agrawal , Sushmita Nair

Text-to-image models are becoming increasingly popular, revolutionizing the landscape of digital art creation by enabling highly detailed and creative visual content generation. These models have been widely employed across various domains,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Matthew Zheng , Enis Simsar , Hidir Yesiltepe , Federico Tombari , Joel Simon , Pinar Yanardag

Recent text-to-image (T2I) models have demonstrated impressive capabilities in photorealistic synthesis and instruction following. However, their reliability in knowledge-intensive settings remains largely unexplored. Unlike natural image…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Ran Zhao , Sheng Jin , Size Wu , Kang Liao , Zerui Gong , Zujin Guo , Yang Xiao , Wei Li

This study introduces a dataset consisting of approximately 9,000 images of mechanical mechanisms and their corresponding descriptions, aimed at supporting research in mechanism design. The dataset consists of a diverse collection of 2D and…

Computer Vision and Pattern Recognition · Computer Science 2024-09-09 Farshid Ghezelbash , Amir Hossein Eskandari , Amir J Bidhendi

Multimodal Information Retrieval has made significant progress in recent years, leveraging the increasingly strong multimodal abilities of deep pre-trained models to represent information across modalities. Music Information Retrieval…

Information Retrieval · Computer Science 2026-02-13 Benjamin Clavié , Atoof Shakir , Jonah Turner , Sean Lee , Aamir Shakir , Makoto P. Kato

Recent advancements in Text-to-Song generation have enabled realistic musical content production, yet existing evaluation benchmarks lack the professional granularity to capture multi-dimensional aesthetic nuances. In this paper, we propose…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-30 Dapeng Wu , Shun Lei , Wei Tan , Guangzheng Li , Yunzhe Wang , Huaicheng Zhang , Lishi Zuo , Zhiyong Wu

Differentially private (DP) image synthesis aims to generate artificial images that retain the properties of sensitive images while protecting the privacy of individual images within the dataset. Despite recent advancements, we find that…

Cryptography and Security · Computer Science 2025-09-01 Chen Gong , Kecen Li , Zinan Lin , Tianhao Wang

Human annotators typically provide annotated data for training machine learning models, such as neural networks. Yet, human annotations are subject to noise, impairing generalization performances. Methodological research on approaches…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Marek Herde , Denis Huseljic , Lukas Rauch , Bernhard Sick

One of the key challenges of detecting AI-generated images is spotting images that have been created by previously unseen generative models. We argue that the limited diversity of the training data is a major obstacle to addressing this…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Jeongsoo Park , Andrew Owens

Improving visual text synthesis has long been a challenging and evolving frontier for image generation models. While recent state-of-the-art (SOTA) models have made remarkable strides in text generation capabilities, existing benchmarks…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Peirong Zhang , Haowei Xu , Jiaxin Zhang , Xuhan Zheng , Guitao Xu , Yuyi Zhang , Junle Liu , Zhenhua Yang , Wei Zhou , Lianwen Jin

While traditional tree-based ensemble methods have long dominated tabular tasks, deep neural networks and emerging foundation models have challenged this primacy, yet no consensus exists on a universally superior paradigm. Existing…

Imbalanced learning remains a fundamental challenge in tabular data applications. Despite decades of research and numerous proposed algorithms, a systematic empirical understanding of how different imbalanced learning methods behave across…

Machine Learning · Computer Science 2026-05-15 Ruizhe Liu , Jiaqi Luo

Can we derive computational metrics to quantify visual creativity in drawings across intelligent agents, while accounting for inherent differences in technical skill and style? To answer this, we curate a novel dataset consisting of 1338…

Human-Computer Interaction · Computer Science 2025-02-11 Surabhi S Nath , Guiomar del Cuvillo y Schröder , Claire E. Stevenson

Recent advancements in image editing have enabled highly controllable and semantically-aware alteration of visual content, posing unprecedented challenges to manipulation localization. However, existing AI-generated forgery localization…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Shiyu Wu , Shuyan Li , Jing Li , Jing Liu , Yequan Wang