Are Language Models Models?
2026/01/15 by Philip Resnik · 1 voice · 1 citation
Computer Science · #cs.CL #cs.AI
paper · pdf
Abstract
Futrell and Mahowald claim LMs "serve as model systems", but an assessment at each of Marr's three levels suggests the claim is clearly not true at the implementation level, poorly motivated at the algorithmic-representational level, and problematic at the computational theory level. LMs are good candidates as tools; calling them cognitive models overstates the case and unnecessarily feeds LLM hype.
Citations
- What enables human language? A biocultural framework
- Dialogue Is Not Enough to Make a Communicative BabyLM (But Neither Is Developmentally Inspired Reinforcement Learning)
- Teacher Demonstrations in a BabyLM's Zone of Proximal Development for Contingent Multi-Turn Interaction
- What Can String Probability Tell Us About Grammaticality?
- Are BabyLMs Deaf to Gricean Maxims? A Pragmatic Evaluation of Sample-efficient Language Models
- Emergent morpho-phonological representations in self-supervised speech models
- Once Upon a Time: Interactive Learning for Storytelling with Small Language Models
- How Causal Abstraction Underpins Computational Explanation
- ROSE: A Universal Neural Grammar
- The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
- What Has a Foundation Model Found? Using Inductive Bias to Probe for World Models
- Information Locality as an Inductive Bias for Neural Language Models
- Constructing language: a framework for explaining acquisition
- Causal Interventions Reveal Shared Structure Across English Filler-Gap Constructions
- On The Landscape of Spoken Language Models: A Comprehensive Survey
- Both Direct and Indirect Evidence Contribute to Dative Alternation Preferences in Language Models
- LLMs syntactically adapt their language use to their conversational partner
- Deep Learning is Not So Mysterious or Different
- Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMs
- What’s Surprising About Surprisal
- Can Language Models Learn Typologically Implausible Languages?
- LLMs as a synthesis between symbolic and distributed approaches to language
- Large Language Models as Proxies for Theories of Human Linguistic Cognition
- Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages
- Stochastic Remembering and Distributed Mnemonic Agency
- The metaphors of artificial intelligence
- Why ‘open’ AI systems are actually closed, and why this matters
- Generalizations across filler-gap dependencies in neural language models
- A Psycholinguistic Evaluation of Language Models' Sensitivity to Argument Roles
- Kallini et al. (2024) do not compare impossible languages with constituency-based ones
- Reclaiming AI as a Theoretical Tool for Cognitive Science
- Larger and more instructable language models become less reliable
- Generalized Measures of Anticipation and Responsivity in Online Language Processing
- How do linguistic illusions arise? Rational inference and good-enough processing as competing latent processes within individuals
- The Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation Analysis
- Lexical Surprisal Shapes the Time Course of Syntactic Structure Building
- Human-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0
- Decomposing modal thought.
- AI models collapse when trained on recursively generated data
- Large Language Models are Biased Because They Are Large Language Models
- Code Pretraining Improves Entity Tracking Abilities of Language Models
- Towards a theory of how the structure of language is acquired by deep neural networks
- Filtered Corpus Training (FiCT) Shows that Language Models can Generalize from Indirect Evidence
- Filtered Corpus Training (FiCT) Shows that Language Models Can Generalize from Indirect Evidence
- The Platonic Representation Hypothesis
- People cannot distinguish GPT-4 from a human in a Turing test
- Language in Vivo vs. in Silico: Size Matters but Larger Language Models Still Do Not Comprehend Language on a Par with Humans Due to Impenetrable Semantic Reference
- Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs
- Encoding of lexical tone in self-supervised models of spoken language
- Bridging the Empirical-Theoretical Gap in Neural Network Formal Language Learning Using Minimum Description Length
- Why are Sensitive Functions Hard for Transformers?
- Learning Syntax Without Planting Trees: Understanding Hierarchical Generalization in Transformers
- Grounded language acquisition through the eyes and ears of a single child
- Investigating Critical Period Effects in Language Acquisition through Neural Language Models
- Mission: Impossible Language Models
- Power Hungry Processing: Watts Driving the Cost of AI Deployment?
- A U-turn on Double Descent: Rethinking Parameter Counting in Statistical Learning
- Subspace Chronicles: How Linguistic Information Emerges, Shifts and Interacts during Language Model Training
- Information Value: Measuring Utterance Predictability as Distance from Plausible Alternatives
- Characterizing Learning Curves During Language Model Pre-Training: Learning, Forgetting, and Stability
- A Sentence is Worth a Thousand Pictures: Can Large Language Models Understand Hum4n L4ngu4ge and the W0rld behind W0rds?
- Opening up ChatGPT: Tracking openness, transparency, and accountability in instruction-tuned text generators
- Testing the Predictions of Surprisal Theory in 11 Languages
- What Do Self-Supervised Speech Models Know About Words?
- Probing self-supervised speech models for phonetic and phonemic information: a case study in aspiration
- Mathematical Structure of Syntactic Merge
- Modeling rapid language learning by distilling Bayesian priors into artificial neural networks
- Physics of Language Models: Part 1, Learning Hierarchical Language Structures
- Putting Natural in Natural Language Processing
- Why open-source generative AI models are an ethical way forward for science
- ROSE: A Neurocomputational Architecture for Syntax
- Testing AI on language comprehension tasks reveals insensitivity to underlying meaning
- ProsAudit, a prosodic benchmark for self-supervised speech models
- How poor is the stimulus? Evaluating hierarchical generalization in neural networks trained on child-directed speech
- Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
- Why Does Surprisal From Larger Transformer-Based Language Models Provide a Poorer Fit to Human Reading Times?
- Training Trajectories of Language Models Across Scales
- A fine-grained comparison of pragmatic language understanding in humans and language models
- Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model
- Modeling structure-building in the brain with CCG parsing and large language models
- The debate over understanding in AI’s large language models
- Over-reliance on English hinders cognitive science
- Neurocompositional computing: From the Central Paradox of Cognition to a new generation of AI systems
- How much do language models copy from their training data? Evaluating linguistic novelty in text generation using RAVEN
- Minimum Description Length Recurrent Neural Networks
- Systematic Inequalities in Language Technology Performance across the World's Languages
- On Logical Inference over Brains, Behaviour, and Artificial Neural Networks
- Word Acquisition in Neural Language Models
- Deduplicating Training Data Makes Language Models Better
- Sensitivity as a Complexity Measure for Sequence Classification Tasks
- Understanding deep learning (still) requires rethinking generalization
- The Theory Crisis in Psychology: How to Move Forward
- Deep Subjecthood: Higher-Order Grammatical Features in Multilingual BERT
- Inductive Biases for Deep Learning of Higher-Level Cognition
- On the Practical Ability of Recurrent Neural Networks to Recognize\n Hierarchical Languages
- RNNs can generate bounded hierarchical languages with optimal memory
- Word predictability effects are linear, not logarithmic: Implications for probabilistic models of sentence comprehension
- Universal linguistic inductive biases via meta-learning
- The State and Fate of Linguistic Diversity and Inclusion in the NLP\n World
- Recurrent Neural Network Language Models Always Learn English-Like Relative Clause Attachment
- BLiMP: The Benchmark of Linguistic Minimal Pairs for English
- Theoretical Limitations of Self-Attention in Neural Sequence Models
- On the Computational Power of RNNs
- What happened to cognitive science?
- Neural Language Models as Psycholinguistic Subjects: Representations of\n Syntactic State
- What do RNN Language Models Learn about Filler-Gap Dependencies?
- On the Practical Computational Power of Finite Precision RNNs for Language Recognition
- Revisiting the poverty of the stimulus: hierarchical generalization without a hierarchical bias in recurrent neural networks
- Assessing the Ability of LSTMs to Learn Syntax-Sensitive Dependencies
- Why Neurons Have Thousands of Synapses, a Theory of Sequence Memory in Neocortex
- Mathematical Foundations for a Compositional Distributional Model of Meaning
- Expectation-based syntactic comprehension
- Finding Structure in Time
- Self-organized formation of topologically correct feature maps
- Critical period effects in second language learning: The influence of maturational state on the acquisition of English as a second language
- The TRACE model of speech perception
Cited by
Discussions
Related