中文
相关论文

相关论文: A very preliminary analysis of DALL-E 2

200 篇论文

In this study we compared how well DALL-E 2 visually represented the meaning of linguistic prompts also given to young children in comprehension tests. Sentences representing fundamental components of grammatical knowledge were selected…

计算与语言 · 计算机科学 2024-03-20 Elliot Murphy , Jill de Villiers , Sofia Lucero Morales

Machine intelligence is increasingly being linked to claims about sentience, language processing, and an ability to comprehend and transform natural language into a range of stimuli. We systematically analyze the ability of DALL-E 2 to…

计算与语言 · 计算机科学 2022-10-26 Evelina Leivada , Elliot Murphy , Gary Marcus

We conduct a pilot study selectively evaluating the cognitive abilities (decision making and spatial reasoning) of two recently released generative transformer models, ChatGPT and DALL-E 2. Input prompts were constructed following neutral a…

人工智能 · 计算机科学 2023-02-21 Zhisheng Tang , Mayank Kejriwal

Whereas generative adversarial networks are capable of synthesizing highly realistic images of faces, cats, landscapes, or almost any other single category, paint-by-text synthesis engines can -- from a single text prompt -- synthesize…

计算机视觉与模式识别 · 计算机科学 2022-08-02 Hany Farid

Type "a sea otter with a pearl earring by Johannes Vermeer" or "a photo of a teddy bear on a skateboard in Times Square" into OpenAI's DALL-E-2 paint-by-text synthesis engine and you will not be disappointed by the delightful and eerily…

图形学 · 计算机科学 2022-06-30 Hany Farid

The field of multimodal research focusing on the comprehension and creation of both images and text has witnessed significant strides. This progress is exemplified by the emergence of sophisticated models dedicated to image captioning at…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Hang Li , Jindong Gu , Rajat Koner , Sahand Sharifzadeh , Volker Tresp

Recent work has shown that despite their impressive capabilities, text-to-image diffusion models such as DALL-E 2 (Ramesh et al., 2022) can display strange behaviours when a prompt contains a word with multiple possible meanings, often…

计算与语言 · 计算机科学 2022-11-24 Jennifer C. White , Ryan Cotterell

We study the way DALLE-2 maps symbols (words) in the prompt to their references (entities or properties of entities in the generated image). We show that in stark contrast to the way human process language, DALLE-2 does not follow the…

计算与语言 · 计算机科学 2022-10-20 Royi Rassin , Shauli Ravfogel , Yoav Goldberg

Recent advancements in language-image models have led to the development of highly realistic images that can be generated from textual descriptions. However, the increased visual quality of these generated images poses a potential threat to…

计算机视觉与模式识别 · 计算机科学 2023-04-17 Shan Jia , Mingzhen Huang , Zhou Zhou , Yan Ju , Jialing Cai , Siwei Lyu

We provide a new multi-task benchmark for evaluating text-to-image models. We perform a human evaluation comparing the most common open-source (Stable Diffusion) and commercial (DALL-E 2) models. Twenty computer science AI graduate students…

With the rapid development of Artificial Intelligence Generated Content (AIGC), it has become a common practice to train models on synthetic data due to data-scarcity and privacy leakage problems. Owing to massive and diverse information…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Shiye Lei , Hao Chen , Sen Zhang , Bo Zhao , Dacheng Tao

Text-to-image AI are capable of generating novel images for inspiration, but their applications for 3D design workflows and how designers can build 3D models using AI-provided inspiration have not yet been explored. To investigate this, we…

人机交互 · 计算机科学 2023-08-02 Vivian Liu , Jo Vermeulen , George Fitzmaurice , Justin Matejka

While recent advancements in artificial intelligence (AI) language models demonstrate cutting-edge performance when working with English texts, equivalent models do not exist in other languages or do not reach the same performance level.…

计算与语言 · 计算机科学 2022-12-26 Noga Mudrik , Adam S. Charles

We propose a new paradigm to automatically generate training data with accurate labels at scale using the text-toimage synthesis frameworks (e.g., DALL-E, Stable Diffusion, etc.). The proposed approach decouples training data generation…

计算机视觉与模式识别 · 计算机科学 2022-12-23 Yunhao Ge , Jiashu Xu , Brian Nlong Zhao , Neel Joshi , Laurent Itti , Vibhav Vineet

Image2Speech is the relatively new task of generating a spoken description of an image. This paper presents an investigation into the evaluation of this task. For this, first an Image2Speech system was implemented which generates image…

计算与语言 · 计算机科学 2020-08-03 Justin van der Hout , Zoltán D'Haese , Mark Hasegawa-Johnson , Odette Scharenborg

Visual language reasoning requires a system to extract text or numbers from information-dense images like charts or plots and perform logical or arithmetic reasoning to arrive at an answer. To tackle this task, existing work relies on…

计算与语言 · 计算机科学 2023-10-05 Peifang Wang , Olga Golovneva , Armen Aghajanyan , Xiang Ren , Muhao Chen , Asli Celikyilmaz , Maryam Fazel-Zarandi

The image captioning task is about to generate suitable descriptions from images. For this task there can be several challenges such as accuracy, fluency and diversity. However there are few metrics that can cover all these properties while…

计算机视觉与模式识别 · 计算机科学 2020-12-15 Chao Zeng , Sam Kwong

Generative AI systems are increasingly capable of expressing emotions via text and imagery. Effective emotional expression will likely play a major role in the efficacy of AI systems -- particularly those designed to support human mental…

Generative AI models like DALL-E 2 can interpret textual prompts and generate high-quality images exhibiting human creativity. Though public enthusiasm is booming, systematic auditing of potential gender biases in AI-generated images…

计算机视觉与模式识别 · 计算机科学 2024-09-13 Luhang Sun , Mian Wei , Yibing Sun , Yoo Ji Suh , Liwei Shen , Sijia Yang

Answering visual questions need acquire daily common knowledge and model the semantic connection among different parts in images, which is too difficult for VQA systems to learn from images with the only supervision from answers. Meanwhile,…

计算与语言 · 计算机科学 2018-05-23 Jialin Wu , Zeyuan Hu , Raymond J. Mooney
‹ 上一页 1 2 3 10 下一页 ›