vix.ing · top · new · best · stats

Words That Make Language Models Perceive

2025/10/02 by S. Q. Wang, Sophie L. Wang, Phillip Isola +4 · 4 voices · 4 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Generative Adversarial Networks and Image Synthesis #Language model #Multimodal Machine Learning Applications #Perception #Sensory cue #Sensory system #Test (biology) #Visual perception

paper · pdf · doi:10.48550/arxiv.2510.02425

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2025/10/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

Large language models (LLMs) trained purely on text ostensibly lack any direct perceptual experience, yet their internal representations are implicitly shaped by multimodal regularities encoded in language. We test the hypothesis that explicit sensory prompting can surface this latent structure, bringing a text-only LLM into closer representational alignment with specialist vision and audio encoders. When a sensory prompt tells the model to 'see' or 'hear', it cues the model to resolve its next-token predictions as if they were conditioned on latent visual or auditory evidence that is never actually supplied. Our findings reveal that lightweight prompt engineering can reliably activate modality-appropriate representations in purely text-trained LLMs.

Citations

Cited by

Discussions

Related