vix.ing · top · new · best · stats · spec

Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics

2025/09/12 by Ryan, Yuriel, Rui Tan, Tan, Rui Yang +4 · 2 citations
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Games #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Humor Studies and Applications #Multimodal Machine Learning Applications

paper · pdf · doi:10.48550/arxiv.2509.12248

openalex publication_date 2025/09/12 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Understanding humor is a core aspect of social intelligence, yet it remains a significant challenge for Large Multimodal Models (LMMs). We introduce PixelHumor, a benchmark dataset of 2,800 annotated multi-panel comics designed to evaluate LMMs' ability to interpret multimodal humor and recognize narrative sequences. Experiments with state-of-the-art LMMs reveal substantial gaps: for instance, top models achieve only 61% accuracy in panel sequencing, far below human performance. This underscores critical limitations in current models' integration of visual and textual cues for coherent narrative and humor understanding. By providing a rigorous framework for evaluating multimodal contextual and narrative reasoning, PixelHumor aims to drive the development of LMMs that better engage in natural, socially aware interactions.

Citations

Cited by

Related