vix.ing · top · new · best · stats · spec

Text Is Not All You Need: Multimodal Prompting Helps LLMs Understand Humor

2024/12/01 by Ashwin Baluja, Baluja, Ashwin · 4 citations
Arts and Humanities · Psychology · #American Literature and Humor Studies #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Humor Studies and Applications #Language, Metaphor, and Cognition

paper · pdf · doi:10.48550/arxiv.2412.05315

openalex publication_date 2024/12/01 · openalex created_date 2024/12/12 · openalex updated_date 2026/07/28

Abstract

While Large Language Models (LLMs) have demonstrated impressive natural language understanding capabilities across various text-based tasks, understanding humor has remained a persistent challenge. Humor is frequently multimodal, relying on phonetic ambiguity, rhythm and timing to convey meaning. In this study, we explore a simple multimodal prompting approach to humor understanding and explanation. We present an LLM with both the text and the spoken form of a joke, generated using an off-the-shelf text-to-speech (TTS) system. Using multimodal cues improves the explanations of humor compared to textual prompts across all tested datasets.

Cited by

Related