中文
相关论文

相关论文: Cracking the Code of Juxtaposition: Can AI Models …

200 篇论文

Understanding humor-particularly when it involves complex, contradictory narratives that require comparative reasoning-remains a significant challenge for large vision-language models (VLMs). This limitation hinders AI's ability to engage…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Tuo Liang , Zhe Hu , Jing Li , Hao Zhang , Yiren Lu , Yunlai Zhou , Yiran Qiao , Disheng Liu , Jeirui Peng , Jing Ma , Yu Yin

Large neural networks can now generate jokes, but do they really "understand" humor? We challenge AI models with three tasks derived from the New Yorker Cartoon Caption Contest: matching a joke to a cartoon, identifying a winning caption,…

计算与语言 · 计算机科学 2023-07-07 Jack Hessel , Ana Marasović , Jena D. Hwang , Lillian Lee , Jeff Da , Rowan Zellers , Robert Mankoff , Yejin Choi

Despite significant advancements in image segmentation and object detection, understanding complex scenes remains a significant challenge. Here, we focus on graphical humor as a paradigmatic example of image interpretation that requires…

计算机视觉与模式识别 · 计算机科学 2025-10-20 Vedaant Jain , Felipe dos Santos Alves Feitosa , Gabriel Kreiman

Understanding satire and humor is a challenging task for even current Vision-Language models. In this paper, we propose the challenging tasks of Satirical Image Detection (detecting whether an image is satirical), Understanding (generating…

计算机视觉与模式识别 · 计算机科学 2024-09-23 Abhilash Nandy , Yash Agarwal , Ashish Patwa , Millon Madhur Das , Aman Bansal , Ankit Raj , Pawan Goyal , Niloy Ganguly

Understanding humor is a core aspect of social intelligence, yet it remains a significant challenge for Large Multimodal Models (LMMs). We introduce PixelHumor, a benchmark dataset of 2,800 annotated multi-panel comics designed to evaluate…

计算机视觉与模式识别 · 计算机科学 2025-09-18 Yuriel Ryan , Rui Yang Tan , Kenny Tsu Wei Choo , Roy Ka-Wei Lee

We present HumorBench, a benchmark designed to evaluate large language models' (LLMs) ability to reason about and explain sophisticated humor in cartoon captions. As reasoning models increasingly saturate existing benchmarks in mathematics…

Large Language Models (LLMs) have shown significant limitations in understanding creative content, as demonstrated by Hessel et al. (2023)'s influential work on the New Yorker Cartoon Caption Contest (NYCCC). Their study exposed a…

Humor recognition has been widely studied as a text classification problem using data-driven approaches. However, most existing work does not examine the actual joke mechanism to understand humor. We break down any joke into two distinct…

计算与语言 · 计算机科学 2021-08-11 Yubo Xie , Junze Li , Pearl Pu

Humor is a central aspect of human communication that has not been solved for artificial agents so far. Large language models (LLMs) are increasingly able to capture implicit and contextual information. Especially, OpenAI's ChatGPT recently…

人工智能 · 计算机科学 2023-06-08 Sophie Jentzsch , Kristian Kersting

Most language models currently available are prone to self-contradiction during dialogues. To mitigate this issue, this study explores a novel contradictory dialogue processing task that aims to detect and modify contradictory statements in…

计算与语言 · 计算机科学 2024-10-08 Xiaofei Wen , Bangzheng Li , Tenghao Huang , Muhao Chen

Comedy serves as a profound reflection of the times we live in and is a staple element of human interactions. In light of the widespread adoption of Large Language Models (LLMs), the intersection of humor and AI has become no laughing…

计算与语言 · 计算机科学 2025-11-25 Adrianna Romanowski , Pedro H. V. Valois , Kazuhiro Fukui

Humans possess multimodal literacy, allowing them to actively integrate information from various modalities to form reasoning. Faced with challenges like lexical ambiguity in text, we supplement this with other modalities, such as thumbnail…

计算机视觉与模式识别 · 计算机科学 2024-10-24 Jiwan Chung , Seungwon Lim , Jaehyun Jeon , Seungbeen Lee , Youngjae Yu

Humor is one of the few cognitive tasks where getting the reasoning right matters as much as getting the answer right. While recent work evaluates humor understanding on benchmarks such as the New Yorker Cartoon Caption Contest (NYCC), it…

人工智能 · 计算机科学 2026-04-17 Hatice Merve Vural , Doga Kukul , Ege Erdem Ozlu , Demir Ekin Arikan , Bob Mankoff , Erkut Erdem , Aykut Erdem

Humor holds up a mirror to social perception: what we find funny often reflects who we are and how we judge others. When language models engage with humor, their reactions expose the social assumptions they have internalized from training…

计算与语言 · 计算机科学 2026-04-22 Shubin Kim , Yejin Son , Junyeong Park , Keummin Ka , Seungbeen Lee , Jaeyoung Lee , Hyeju Jang , Alice Oh , Youngjae Yu

Multimodal Large Language Models are primarily trained and evaluated on aligned image-text pairs, which leaves their ability to detect and resolve real-world inconsistencies largely unexplored. In open-domain applications visual and textual…

Humor is a natural and fundamental component of human interactions. When correctly applied, humor allows us to express thoughts and feelings conveniently and effectively, increasing interpersonal affection, likeability, and trust. However,…

计算与语言 · 计算机科学 2020-11-25 Felipe Godoy

We present a novel multimodal preference dataset for creative tasks, consisting of over 250 million human ratings on more than 2.2 million captions, collected through crowdsourcing rating data for The New Yorker's weekly cartoon caption…

With the recent advances in Artificial Intelligence (AI) and Large Language Models (LLMs), the automation of daily tasks, like automatic writing, is getting more and more attention. Hence, efforts have focused on aligning LLMs with human…

计算与语言 · 计算机科学 2025-06-09 Mohammadamin Shafiei , Hamidreza Saffari

Recent multimodal large language models have shown promising ability in generating humorous captions for images, yet they still lack stable control over explicit cultural context, making it difficult to jointly maintain image relevance,…

计算与语言 · 计算机科学 2026-04-21 Run Xu , Lu Li , Rongzhao Zhang , Jie Xu

Vision-language models have recently evolved into versatile systems capable of high performance across a range of tasks, such as document understanding, visual question answering, and grounding, often in zero-shot settings. Comics…

计算机视觉与模式识别 · 计算机科学 2025-01-09 Emanuele Vivoli , Mohamed Ali Souibgui , Andrey Barsky , Artemis LLabrés , Marco Bertini , Dimosthenis Karatzas
‹ 上一页 1 2 3 10 下一页 ›