2026/01/01 by Onur Keleş, Asli Ozyurek, Gerardo Ortega +2 · 1 voice
Computer Science · Psychology · #Categorization, perception, and language #Hand Gesture Recognition Systems #Hearing Impairment and Communication
paper · pdf · doi:10.18653/v1/2026.acl-long.1907
openalex publication_date 2026/01/01 · openalex created_date 2026/07/02 · openalex updated_date 2026/07/29
Iconicity, the resemblance between linguistic form and meaning, is pervasive in sign languages, offering a natural testbed for visual grounding in vision-language models (VLMs).We introduce the Visual Iconicity Challenge, a video-based benchmark that adapts psycholinguistic measures to evaluate VLMs on three tasks: (i) phonological sign-form prediction, (ii) transparency (inferring meaning from visual form), and (iii) graded iconicity ratings.We assess 17 state-of-the-art VLMs in zeroand few-shot settings on Sign Language of the Netherlands and compare them to human baselines.VLMs mirror human phonological difficulty patterns (e.g., handshape harder than location) and achieve moderate to strong alignment with human iconicity ratings.However, most of them still fail to infer lexical meaning from visual form alone and show a systematic objectbased bias that inverts the human preference for action-based signs.Crucially, models with stronger phonological form prediction correlate better with human iconicity judgments, indicating shared sensitivity to visually grounded structure.Our findings validate these diagnostic tasks, show that explicit reasoning narrows the open-to-closed-model calibration gap, and motivate human-centric signals for modelling iconicity in multimodal models.