中文
相关论文

相关论文: Discovering the Hidden Vocabulary of DALLE-2

200 篇论文

Recent work has shown that despite their impressive capabilities, text-to-image diffusion models such as DALL-E 2 (Ramesh et al., 2022) can display strange behaviours when a prompt contains a word with multiple possible meanings, often…

计算与语言 · 计算机科学 2022-11-24 Jennifer C. White , Ryan Cotterell

We study the way DALLE-2 maps symbols (words) in the prompt to their references (entities or properties of entities in the generated image). We show that in stark contrast to the way human process language, DALLE-2 does not follow the…

计算与语言 · 计算机科学 2022-10-20 Royi Rassin , Shauli Ravfogel , Yoav Goldberg

We report a peculiar observation that LLMs can assign hidden meanings to sequences that seem visually incomprehensible to humans: for example, a nonsensical phrase consisting of Byzantine musical symbols is recognized by gpt-4o as "say…

计算与语言 · 计算机科学 2025-03-04 K. O. T. Erziev

The DALL-E 2 system generates original synthetic images corresponding to an input text as caption. We report here on the outcome of fourteen tests of this system designed to assess its common sense, reasoning and ability to understand…

计算机视觉与模式识别 · 计算机科学 2022-05-04 Gary Marcus , Ernest Davis , Scott Aaronson

We discover that many natural-language prompts can be replaced by corresponding prompts that are unintelligible to humans but that provably elicit similar behavior in language models. We call these prompts "evil twins" because they are…

计算与语言 · 计算机科学 2024-10-08 Rimon Melamed , Lucas H. McCabe , Tanay Wakhare , Yejin Kim , H. Howie Huang , Enric Boix-Adsera

In this study we compared how well DALL-E 2 visually represented the meaning of linguistic prompts also given to young children in comprehension tests. Sentences representing fundamental components of grammatical knowledge were selected…

计算与语言 · 计算机科学 2024-03-20 Elliot Murphy , Jill de Villiers , Sofia Lucero Morales

A word may contain one or more hidden concepts. While the "animal" word evokes many images in our minds and encapsulates many concepts (birds, dogs, cats, crocodiles, etc.), the `parrot' word evokes a single image (a colored bird with a…

计算与语言 · 计算机科学 2023-12-29 İlknur Dönmez Phd , Mehmet Haklıdır Phd

Machine intelligence is increasingly being linked to claims about sentience, language processing, and an ability to comprehend and transform natural language into a range of stimuli. We systematically analyze the ability of DALL-E 2 to…

计算与语言 · 计算机科学 2022-10-26 Evelina Leivada , Elliot Murphy , Gary Marcus

Type "a sea otter with a pearl earring by Johannes Vermeer" or "a photo of a teddy bear on a skateboard in Times Square" into OpenAI's DALL-E-2 paint-by-text synthesis engine and you will not be disappointed by the delightful and eerily…

图形学 · 计算机科学 2022-06-30 Hany Farid

Probabilistic puzzles can be confusing, partly because they are formulated in natural languages - full of unclarities and ambiguities - and partly because there is no widely accepted and intuitive formal language to express them. We propose…

计算机科学中的逻辑 · 计算机科学 2025-04-11 Elena Di Lavore , Bart Jacobs , Mario Román

Word clouds became a standard tool for presenting results of natural language processing methods such as topic modelling. They exhibit most important words, where word size is often chosen proportional to the relevance of words within a…

统计计算 · 统计学 2023-02-14 Peter Winker

Riddles are concise linguistic puzzles that describe an object or idea through indirect, figurative, or playful clues. They are a longstanding form of creative expression, requiring the solver to interpret hints, recognize patterns, and…

计算与语言 · 计算机科学 2026-01-28 Niharika Sri Parasa , Chaitali Diwan , Srinath Srinivasa

Open-vocabulary models are a promising new paradigm for image classification. Unlike traditional classification models, open-vocabulary models classify among any arbitrary set of categories specified with natural language during inference.…

计算机视觉与模式识别 · 计算机科学 2023-12-05 Sarah Pratt , Ian Covert , Rosanne Liu , Ali Farhadi

Dogwhistles are coded expressions that simultaneously convey one meaning to a broad audience and a second one, often hateful or provocative, to a narrow in-group; they are deployed to evade both political repercussions and algorithmic…

计算与语言 · 计算机科学 2025-02-26 Julia Mendelsohn , Ronan Le Bras , Yejin Choi , Maarten Sap

Text-guided image generation models can be prompted to generate images using nonce words adversarially designed to robustly evoke specific visual concepts. Two approaches for such generation are introduced: macaronic prompting, which…

计算机视觉与模式识别 · 计算机科学 2022-08-09 Raphaël Millière

Detecting ambiguity is important for language understanding, including uncertainty estimation, humour detection, and processing garden path sentences. We assess language models' sensitivity to ambiguity by introducing an adversarial…

计算与语言 · 计算机科学 2025-06-03 Antonia Karamolegkou , Oliver Eberle , Phillip Rust , Carina Kauf , Anders Søgaard

Language models produce a distribution over the next token; can we use this information to recover the prompt tokens? We consider the problem of language model inversion and show that next-token probabilities contain a surprising amount of…

计算与语言 · 计算机科学 2023-11-27 John X. Morris , Wenting Zhao , Justin T. Chiu , Vitaly Shmatikov , Alexander M. Rush

An important part of textual inference is making deductions involving monotonicity, that is, determining whether a given assertion entails restrictions or relaxations of that assertion. For instance, the statement 'We know the epidemic…

计算与语言 · 计算机科学 2009-06-16 Cristian Danescu-Niculescu-Mizil , Lillian Lee , Richard Ducott

Large Vision-Language Models (VLMs) have demonstrated strong capabilities in tasks requiring a fine-grained understanding of literal meaning in images and text, such as visual question-answering or visual entailment. However, there has been…

计算与语言 · 计算机科学 2025-02-18 Arkadiy Saakyan , Shreyas Kulkarni , Tuhin Chakrabarty , Smaranda Muresan

The article describes a model of automatic analysis of puns, where a word is intentionally used in two meanings at the same time (the target word). We employ Roget's Thesaurus to discover two groups of words which, in a pun, form around two…

计算与语言 · 计算机科学 2017-07-19 Elena Mikhalkova , Yuri Karyakin
‹ 上一页 1 2 3 10 下一页 ›