中文
相关论文

相关论文: SoS: Analysis of Surface over Semantics in Multili…

200 篇论文

Text-to-image generation models have achieved strong performance in culturally homogeneous settings, yet their ability to generate multicultural scenes, where people and landmarks originate from different cultures, remains largely…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Parth Bhalerao , Mounika Yalamarty , Brian Trinh , Oana Ignat

Understanding spatial relations is a crucial cognitive ability for both humans and AI. While current research has predominantly focused on the benchmarking of text-to-image (T2I) models, we propose a more comprehensive evaluation that…

计算机视觉与模式识别 · 计算机科学 2024-11-13 Shang Hong Sim , Clarence Lee , Alvin Tan , Cheston Tan

Accurate interpretation and visualization of human instructions are crucial for text-to-image (T2I) synthesis. However, current models struggle to capture semantic variations from word order changes, and existing evaluations, relying on…

计算与语言 · 计算机科学 2025-04-18 Xiangru Zhu , Penglei Sun , Yaoxian Song , Yanghua Xiao , Zhixu Li , Chengyu Wang , Jun Huang , Bei Yang , Xiaoxiao Xu

We propose T2I-ReasonBench, a benchmark evaluating reasoning capabilities of text-to-image (T2I) models. It consists of four dimensions: Idiom Interpretation, Textual Image Design, Entity-Reasoning and Scientific-Reasoning. We propose a…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Kaiyue Sun , Rongyao Fang , Chengqi Duan , Xian Liu , Xihui Liu

The transformative potential of text-to-image (T2I) models hinges on their ability to synthesize culturally diverse, photorealistic images from textual prompts. However, these models often perpetuate cultural biases embedded within their…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Muna Numan Said , Aarib Zaidi , Rabia Usman , Sonia Okon , Praneeth Medepalli , Kevin Zhu , Vasu Sharma , Sean O'Brien

Evidence shows that text-to-image (T2I) models disproportionately reflect Western cultural norms, amplifying misrepresentation and harms to minority groups. However, evaluating cultural sensitivity is inherently complex due to its fluid and…

Synthetic face generation has rapidly advanced with the emergence of text-to-image (T2I) and of multimodal large language models, enabling high-fidelity image production from natural-language prompts. Despite the widespread adoption of…

计算机与社会 · 计算机科学 2026-02-04 Mengting Wei , Aditya Gulati , Guoying Zhao , Nuria Oliver

Subject-consistent generation (SCG)-aiming to maintain a consistent subject identity across diverse scenes-remains a challenge for text-to-image (T2I) models. Existing training-free SCG methods often achieve consistency at the cost of…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Zhanxin Gao , Beier Zhu , Liang Yao , Jian Yang , Ying Tai

Text-to-Image (T2I) models have recently gained significant attention due to their ability to generate high-quality images and are consequently used in a wide range of applications. However, there are concerns about the gender bias of these…

计算机视觉与模式识别 · 计算机科学 2025-07-28 Yunbo Lyu , Zhou Yang , Yuqing Niu , Jing Jiang , David Lo

Text-to-image (T2I) models offer great potential for creating virtually limitless synthetic data, a valuable resource compared to fixed and finite real datasets. Previous works evaluate the utility of synthetic data from T2I models on three…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Zhang Xiaofeng , Aaron Courville , Michal Drozdzal , Adriana Romero-Soriano

Text-to-Image (T2I) models are being increasingly adopted in diverse global communities where they create visual representations of their unique cultures. Current T2I benchmarks primarily focus on faithfulness, aesthetics, and realism of…

计算机视觉与模式识别 · 计算机科学 2025-01-22 Nithish Kannen , Arif Ahmad , Marco Andreetto , Vinodkumar Prabhakaran , Utsav Prabhu , Adji Bousso Dieng , Pushpak Bhattacharyya , Shachi Dave

We provide a new multi-task benchmark for evaluating text-to-image models. We perform a human evaluation comparing the most common open-source (Stable Diffusion) and commercial (DALL-E 2) models. Twenty computer science AI graduate students…

The rapid advancements of Text-to-Image (T2I) models have ushered in a new phase of AI-generated content, marked by their growing ability to interpret and follow user instructions. However, existing T2I model evaluation benchmarks fall…

计算机视觉与模式识别 · 计算机科学 2025-06-26 Xinyu Wei , Jinrui Zhang , Zeqing Wang , Hongyang Wei , Zhen Guo , Lei Zhang

User prompts for generative AI models are often underspecified, leading to a misalignment between the user intent and models' understanding. As a result, users commonly have to painstakingly refine their prompts. We study this alignment…

人工智能 · 计算机科学 2025-10-27 Meera Hahn , Wenjun Zeng , Nithish Kannen , Rich Galt , Kartikeya Badola , Been Kim , Zi Wang

Text-to-Image (TTI) models are powerful creative tools but risk amplifying harmful social biases. We frame representational societal bias assessment as an image curation and evaluation task and introduce a pilot benchmark of occupational…

计算与语言 · 计算机科学 2025-09-03 Shaina Raza , Maximus Powers , Partha Pratim Saha , Mahveen Raza , Rizwan Qureshi

Image to Image Translation (I2I) is a challenging computer vision problem used in numerous domains for multiple tasks. Recently, ophthalmology became one of the major fields where the application of I2I is increasing rapidly. One such…

图像与视频处理 · 电气工程与系统科学 2021-12-14 Hemanth Pasupuleti , G. N. Girish

Large Text-to-Image(T2I) diffusion models have shown a remarkable capability to produce photorealistic and diverse images based on text inputs. However, existing works only support limited language input, e.g., English, Chinese, and…

计算机视觉与模式识别 · 计算机科学 2023-08-24 Fulong Ye , Guang Liu , Xinya Wu , Ledell Wu

The prosody of a spoken utterance, including features like stress, intonation and rhythm, can significantly affect the underlying semantics, and as a consequence can also affect its textual translation. Nevertheless, prosody is rarely…

计算与语言 · 计算机科学 2024-11-01 Ioannis Tsiamas , Matthias Sperber , Andrew Finch , Sarthak Garg

Text-to-image generative models are a new and powerful way to generate visual artwork. However, the open-ended nature of text as interaction is double-edged; while users can input anything and have access to an infinite range of…

人机交互 · 计算机科学 2023-09-29 Vivian Liu , Lydia B. Chilton

Advanced diffusion-based Text-to-Image (T2I) models, such as the Stable Diffusion Model, have made significant progress in generating diverse and high-quality images using text prompts alone. However, when non-famous users require…

计算机视觉与模式识别 · 计算机科学 2024-03-25 Yang Li , Songlin Yang , Wei Wang , Jing Dong