English

Text Is Not All You Need: Multimodal Prompting Helps LLMs Understand Humor

Computation and Language 2024-12-10 v1 Computers and Society

Abstract

While Large Language Models (LLMs) have demonstrated impressive natural language understanding capabilities across various text-based tasks, understanding humor has remained a persistent challenge. Humor is frequently multimodal, relying on phonetic ambiguity, rhythm and timing to convey meaning. In this study, we explore a simple multimodal prompting approach to humor understanding and explanation. We present an LLM with both the text and the spoken form of a joke, generated using an off-the-shelf text-to-speech (TTS) system. Using multimodal cues improves the explanations of humor compared to textual prompts across all tested datasets.

Keywords

Cite

@article{arxiv.2412.05315,
  title  = {Text Is Not All You Need: Multimodal Prompting Helps LLMs Understand Humor},
  author = {Ashwin Baluja},
  journal= {arXiv preprint arXiv:2412.05315},
  year   = {2024}
}