中文
相关论文

相关论文: What You See is What You Read? Improving Text-Imag…

200 篇论文

Beyond conventional paradigms of translating speech and text, recently, there has been interest in automated transcreation of images to facilitate localization of visual content across different cultures. Attempts to define this as a formal…

计算与语言 · 计算机科学 2025-03-24 Simran Khanuja , Vivek Iyer , Claire He , Graham Neubig

Multilingual alignment of sentence representations has mostly required bitexts to bridge the gap between languages. We investigate whether visual information can bridge this gap instead. Image caption datasets are very easy to create…

计算与语言 · 计算机科学 2025-05-21 Nathaniel Krasner , Nicholas Lanuzo , Antonios Anastasopoulos

Although image captioning models have made significant advancements in recent years, the majority of them heavily depend on high-quality datasets containing paired images and texts which are costly to acquire. Previous works leverage the…

计算机视觉与模式识别 · 计算机科学 2023-12-15 Zhiyue Liu , Jinyuan Liu , Fanrong Ma

In this paper, we introduce a model designed to improve the prediction of image-text alignment, targeting the challenge of compositional understanding in current visual-language models. Our approach focuses on generating high-quality…

计算机视觉与模式识别 · 计算机科学 2024-10-02 Yuheng Li , Haotian Liu , Mu Cai , Yijun Li , Eli Shechtman , Zhe Lin , Yong Jae Lee , Krishna Kumar Singh

Aligning test items to content standards is a critical step in test development to collect validity evidence based on content. Item alignment has typically been conducted by human experts. This judgmental process can be subjective and…

计算与语言 · 计算机科学 2025-10-14 Yanbin Fu , Hong Jiao , Tianyi Zhou , Nan Zhang , Ming Li , Qingshu Xu , Sydney Peters , Robert W. Lissitz

Computational visual aesthetics has recently become an active research area. Existing state-of-art methods formulate this as a binary classification task where a given image is predicted to be beautiful or not. In many applications such as…

计算机视觉与模式识别 · 计算机科学 2017-04-06 Parag S. Chandakkar , Vijetha Gattupalli , Baoxin Li

Language grounded image understanding tasks have often been proposed as a method for evaluating progress in artificial intelligence. Ideally, these tasks should test a plethora of capabilities that integrate computer vision, reasoning, and…

机器学习 · 计算机科学 2019-05-28 Kushal Kafle , Robik Shrestha , Christopher Kanan

This paper introduces a large-scale multimodal and multilingual dataset that aims to facilitate research on grounding words to images in their contextual usage in language. The dataset consists of images selected to unambiguously illustrate…

计算与语言 · 计算机科学 2022-06-20 Josiah Wang , Pranava Madhyastha , Josiel Figueiredo , Chiraag Lala , Lucia Specia

Human beings often assess the aesthetic quality of an image coupled with the identification of the image's semantic content. This paper addresses the correlation issue between automatic aesthetic quality assessment and semantic recognition.…

计算机视觉与模式识别 · 计算机科学 2017-04-05 Yueying Kao , Ran He , Kaiqi Huang

Large-scale vision-language pre-training has shown impressive advances in a wide range of downstream tasks. Existing methods mainly model the cross-modal alignment by the similarity of the global representations of images and texts, or…

计算机视觉与模式识别 · 计算机科学 2022-09-20 Juncheng Li , Xin He , Longhui Wei , Long Qian , Linchao Zhu , Lingxi Xie , Yueting Zhuang , Qi Tian , Siliang Tang

Despite substantial progress in text-to-image generation, achieving precise text-image alignment remains challenging, particularly for prompts with rich compositional structure or imaginative elements. To address this, we introduce Negative…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Sangha Park , Eunji Kim , Yeongtak Oh , Jooyoung Choi , Sungroh Yoon

Textual representations based on pre-trained language models are key, especially in few-shot learning scenarios. What makes a representation good for text classification? Is it due to the geometric properties of the space or because it is…

计算与语言 · 计算机科学 2023-06-01 Cesar Gonzalez-Gutierrez , Audi Primadhanty , Francesco Cazzaro , Ariadna Quattoni

In the quest for fairness in artificial intelligence, novel approaches to enhance it in facial image based gender classification algorithms using text guided methodologies are presented. The core methodology involves leveraging semantic…

计算机视觉与模式识别 · 计算机科学 2025-12-15 Anoop Krishnan

While recent text-to-image models can generate photorealistic images from text prompts that reflect detailed instructions, they still face significant challenges in accurately rendering words in the image. In this paper, we propose to…

计算机视觉与模式识别 · 计算机科学 2025-05-01 Wataru Shimoda , Naoto Inoue , Daichi Haraguchi , Hayato Mitani , Seiichi Uchida , Kota Yamaguchi

Text classification helps analyse texts for semantic meaning and relevance, by mapping the words against this hierarchy. An analysis of various types of texts is invaluable to understanding both their semantic meaning, as well as their…

机器学习 · 计算机科学 2022-11-16 Chaitanya Chadha , Vandit Gupta , Deepak Gupta , Ashish Khanna

Utilizing a shared embedding space, emerging multimodal models exhibit unprecedented zero-shot capabilities. However, the shared embedding space could lead to new vulnerabilities if different modalities can be misaligned. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Shaeke Salman , Md Montasir Bin Shams , Xiuwen Liu

A comprehensive understanding of vision and language and their interrelation are crucial to realize the underlying similarities and differences between these modalities and to learn more generalized, meaningful representations. In recent…

计算机视觉与模式识别 · 计算机科学 2021-12-10 Anindya Sundar Das , Sriparna Saha

When humans read a specific text, they often visualize the corresponding images, and we hope that computers can do the same. Text-to-image synthesis (T2I), which focuses on generating high-quality images from textual descriptions, has…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Nonghai Zhang , Hao Tang

We introduce SeeingSounds, a lightweight and modular framework for audio-to-image generation that leverages the interplay between audio, language, and vision-without requiring any paired audio-visual data or training on visual generative…

People see text. Humans read by recognizing words as visual objects, including their shapes, layouts, and patterns, before connecting them to meaning, which enables us to handle typos, distorted fonts, and various scripts effectively.…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Ling Xing , Rui Yan , Alex Jinpeng Wang , Zechao Li , Jinhui Tang