Probing Classifiers: Promises, Shortcomings, and Advances
2021/02/24 by Yonatan Belinkov, Belinkov, Yonatan · 202 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #cs.CL
paper · pdf · doi:10.48550/arxiv.2102.12452
Accepted to Computational Linguistics as a squib
arxiv created 2021/09/22 · arxiv updated 2021/09/23
Abstract
Probing classifiers have emerged as one of the prominent methodologies for interpreting and analyzing deep neural network models of natural language processing. The basic idea is simple -- a classifier is trained to predict some linguistic property from a model's representations -- and has been used to examine a wide variety of models and properties. However, recent studies have demonstrated various methodological limitations of this approach. This article critically reviews the probing classifiers framework, highlighting their promises, shortcomings, and advances.
Citations
Cited by
- Beyond task performance: Decoding bioacoustic embeddings with speech features
- Two Regimes of Chain-of-Thought Unfaithfulness: Behavioral Detection Fails Where Models Are Wrong
- When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles
- Where meaning lives: Layer-wise accessibility of psycholinguistic features in encoder and decoder language models
- Emergent Latent-State Computation under Stochastic Volatility
- Cheap Probes Predict Expensive Training in 3D-CT Vision--Language Models
- Detect Before You Leap: Mirage Detection in Vision-Language Models
- Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs
- Latent Space Probing for Adult Content Detection in Video Generative Models
- Activations as Features: Probing LLMs for Generalizable Essay Scoring Representations
- In-Context Algebra
- Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants
- The Deleuzian Representation Hypothesis
- Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors
- Behavior and Representation in Large Language Models for Combinatorial Optimization: From Feature Extraction to Algorithm Selection
- Investigating The Functional Roles of Attention Heads in Vision Language Models: Evidence for Reasoning Modules
- Interpreto: An Explainability Library for Transformers
- Mary, the Cheeseburger-Eating Vegetarian: Do LLMs Recognize Incoherence in Narratives?
- Verbalizing LLMs' assumptions to explain and control sycophancy
- Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM Reasoning
- SocialFusion: Addressing Social Degradation in Pre-trained Vision-Language Models
- Freeze, Diffuse, Decode: Geometry-Aware Adaptation of Pretrained Transformer Embeddings for Antimicrobial Peptide Design
- Steering Awareness: Detecting Activation Steering from Within
- Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks
- Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models
- Investigating self-supervised representations for audio-visual deepfake detection
- Cognitive Foundations for Reasoning and Their Manifestation in LLMs
- Towards Understanding Layer Contributions in Tabular In-Context Learning Models
- How Language Directions Align with Token Geometry in Multilingual LLMs
- Feature Quality and Adaptability of Medical Foundation Models: A Comparative Evaluation for Radiographic Classification and Segmentation
- NumPert: Numerical Perturbations to Probe Language Models for Veracity Prediction
- CLIP is All You Need for Human-like Semantic Representations in Stable Diffusion
- MONICA: Real-Time Monitoring and Calibration of Chain-of-Thought Sycophancy in Large Reasoning Models
- Probing the Probes: Methods and Metrics for Concept Alignment
- Interpreting Multi-Attribute Confounding through Numerical Attributes in Large Language Models
- ParaScopes: What do Language Models Activations Encode About Future Text?
- Calibration Across Layers: Understanding Calibration Evolution in LLMs
- Misalignment Has a Personality: A Big Five Account of Emergent Misalignment
- Relation Geometry in Semantic Space of Language Models
- Do BERT Embeddings Encode Narrative Dimensions? A Token-Level Probing Analysis of Time, Space, Causality, and Character in Fiction
- Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks
- Do You Trust Me? Cognitive-Affective Signatures of Trustworthiness in Large Language Models
- A Survey on Unlearning in Large Language Models
- TELL-TALE: Task Efficient LLMs with Task Aware Layer Elimination
- Probing Neural Combinatorial Optimization Models
- Training-Free Spectral Fingerprints of Voice Processing in Transformers
- That's Deprecated! Understanding, Detecting, and Steering Knowledge Conflicts in Language Models for Code Generation
- Context-aware Fairness Evaluation and Mitigation in LLMs
- Semantic Prosody in Machine Translation: the English-Chinese Case of Passive Structures
- CAST: Compositional Analysis via Spectral Tracking for Understanding Transformer Layer Functions
- Probing Latent Knowledge Conflict for Faithful Retrieval-Augmented Generation
- Causality ≠ Decodability, and Vice Versa: Lessons from Interpreting Counting ViTs
- Mapping Semantic & Syntactic Relationships with Geometric Rotation
- Toward Mechanistic Explanation of Deductive Reasoning in Language Models
- On the Alignment Between Supervised and Self-Supervised Contrastive Learning
- Reproducing and Extending Causal Insights Into Term Frequency Computation in Neural Rankers
- GraphGhost: Tracing Structures Behind Large Language Models
- Decoding Emotion in the Deep: A Systematic Study of How LLMs Represent, Retain, and Express Emotion
- Shape Happens: Automatic Feature Manifold Discovery in LLMs via Supervised Multi-Dimensional Scaling
- Analyzing Latent Concepts in Code Language Models
- Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT
- Neural Message-Passing on Attention Graphs for Hallucination Detection
- Emergent World Representations in OpenVLA
- When Can AI Models Explain Learning? Validity Criteria for AI as Cognitive Models in Education
- The Philosophy of Language Models
- Hidden APIs in Language Models: Discovering Reusable Causal Interfaces from Forked Futures
- Interpretable Representation via LLM-Driven Generative Disentanglement for Local-Life Service Recommendation
- STEREODISCO: Discovering Stereotypicality in LLMs
- What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations
- From Found to Designed: Concepts as a Design Axis for Large Language Models
- The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models
- A Pipeline to Assess Merging Methods via Behavior and Internals
- V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models
- AdaptiveK Sparse Autoencoders: Dynamic Sparsity Allocation for Interpretable LLM Representations
- Do Activation Verbalization Methods Convey Privileged Information?
- Steering Language Models in Multi-Token Generation: A Case Study on Tense and Aspect
- Do All Autoregressive Transformers Remember Facts the Same Way? A Cross-Architecture Analysis of Recall Mechanisms
- Interpreting the Effects of Quantization on LLMs
- ALICE: An Interpretable Neural Architecture for Generalization in Substitution Ciphers
- Statistical Methods in Generative AI
- Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
- Not All Splits Are Equal: Rethinking Attribute Generalization Across Unrelated Categories
- A Review of Developmental Interpretability in Large Language Models
- Emphasis Sensitivity in Speech Representations
- Magical: Medical Lay Language Generation via Semantic Invariance and Layperson-tailored Adaptation
- Towards Integrated Alignment
- Balancing Stylization and Truth via Disentangled Representation Steering
- R2-CoD: Understanding Text-Graph Complementarity in Relational Reasoning via Knowledge Co-Distillation
- What Does it Mean for a Neural Network to Learn a "World Model"?
- Length Representations in Large Language Models
- Basic Reading Distillation
- On the Performance of Concept Probing: The Influence of the Data (Extended Version)
- Concept Probing: Where to Find Human-Defined Concepts (Extended Version)
- How does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding
- Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning
- Debiasing Methods in Natural Language Understanding Make Bias More Accessible
- Spectral Principal Paths: A Spectral Perspective on Linear Representation Formation in LLMs
- Using AI to replicate human experimental results: a motion study
- NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment
- The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
- KGRAG-Ex: Explainable Retrieval-Augmented Generation with Knowledge Graph-based Perturbations
- What Has a Foundation Model Found? Using Inductive Bias to Probe for World Models
- LAWFUL: Law-Aligned Witness for Faithful Use of Latents
- The Thin Line Between Comprehension and Persuasion in LLMs
- How Do Vision-Language Models Process Conflicting Information Across Modalities?
- UFM: A Simple Path towards Unified Dense Correspondence with Flow
- Linearly Decoding Refused Knowledge in Aligned Language Models
- Position: Use Sparse Autoencoders to Discover Unknowns
- On the Predictive Power of Representation Dispersion in Language Models
- TuCo: Measuring the Contribution of Fine-Tuning to Individual Responses of LLMs
- Do We Really Need GNNs with Explicit Structural Modeling? MLPs Suffice for Language Model Representations
- Algebraic Priors for Approximately Equivariant Networks
- Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?
- Large Language Models as Psychological Simulators: A Methodological Guide
- Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic Evaluations
- Can structural correspondences ground real world representational content in Large Language Models?
- Addition in Four Movements: Mapping Layer-wise Information Trajectories in LLMs
- RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?
- Enabling Precise Topic Alignment in Large Language Models Via Sparse Autoencoders
- Improving Causal Interventions in Amnesic Probing with Mean Projection or LEACE
- Detecting High-Stakes Interactions with Activation Probes
- Preserving Task-Relevant Information Under Linear Concept Removal
- Spelling-out is not Straightforward: LLMs' Capability of Tokenization from Token to Characters
- System-Aware Unlearning Algorithms: Use Lesser, Forget Faster
- Towards an Explainable Comparison and Alignment of Feature Embeddings
- AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models
- Line of Sight: On Linear Representations in VLLMs
- Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
- Fine-Grained Interpretation of Political Opinions in Large Language Models
- Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey
- Mechanistic Decomposition of Sentence Representations
- Self-Supervised Contrastive Learning is Approximately Supervised Contrastive Learning
- Seeing What Tastes Good: Revisiting Multimodal Distributional Semantics in the Billion Parameter Era
- Linear Spatial World Models Emerge in Large Language Models
- Geospatial Mechanistic Interpretability of Large Language Models
- Quantitative LLM Judges
- Different Speech Translation Models Encode and Translate Speaker Gender Differently
- Probing Neural Topology of Large Language Models
- Knowing Before Saying: LLM Representations Encode Information About Chain-of-Thought Success Before Completion
- Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
- Simulating Early Phonetic and Word Learning Without Linguistic Categories
- From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs
- Response Uncertainty and Probe Modeling: Two Sides of the Same Coin in LLM Interpretability?
- Cultural Awareness is Represented but Not Decoded: Tracing Mythological Knowledge across 18 Open-Source LLMs
- Discriminating Form and Meaning in Multilingual Models with Minimal-Pair ABX Tasks
- Reading Between the Prompts: How Stereotypes Shape LLM's Implicit Personalization
- SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models
- ChaosProbe: A Neurochaotic Lens on Frozen Transformer Input-Embedding Spaces
- Language Models use Lookbacks to Track Beliefs
- Concept Incongruence: An Exploration of Time and Death in Role Playing
- Understanding Task Representations in Neural Networks via Bayesian Ablation
- Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers
- Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models
- Probing the 3D Object-Level Understanding of Pre-Trained Detection Transformers
- Mitigating Content Effects on Reasoning in Language Models through Fine-Grained Activation Steering
- Concept-Guided Interpretability via Neural Chunking
- Probing Subphonemes in Morphology Models
- Designing and Contextualising Probes for African Languages
- Revealing economic facts: LLMs know more than they say
- Interpreting Multilingual and Document-Length Sensitive Relevance Computations in Neural Retrieval Models through Axiomatic Causal Interventions
- Stroke Lesions as a Rosetta Stone for Language Model Interpretability
- Using a Cross-Task Grid of Linear Probes to Interpret CNN Model Predictions On Retinal Images
- Perturbation: A simple and efficient adversarial tracer for representation learning in language models
- Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought
- Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure
- What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness
- Validating Causal Abstraction Metrics on Simulated Complex Systems
- Do Sparse Autoencoders Capture Concept Manifolds?
- Interpreting "Interpretability" and Explaining "Explainability" in Machine Learning in Physics
- Beyond Solving: Prescriptive Probing for Neural Routing Solvers
- Prompt Injection as Role Confusion
- Spilled Energy in Large Language Models
- Bayes-Sufficient Representations in Supervised Learning
- Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning
- Do Language Models Track Entities Across State Changes?
- Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance
- Deep Sequence Modeling with Quantum Dynamics: Language as a Wave Function
- The Truthfulness Spectrum Hypothesis
- Prometheus Mind: Retrofitting Memory to Frozen Language Models
- CLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark
- Reasoning Models Know What's Important, and Encode It in Their Activations
- Why Linear Interpretability Works: Invariant Subspaces as a Result of Architectural Constraints
- Functional Emotions or Situational Contexts? A Discriminating Test from the Mythos Preview System Card
- The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Models
- The Spectral Geometry of Thought: Phase Transitions, Instruction Reversal, Token-Level Dynamics, and Perfect Correctness Prediction in How Transformers Reason
- UniCog: Uncovering Cognitive Abilities of LLMs through Latent Mind Space Analysis
- Probing Character-level Transformers for the Spanish L-shaped Morphome
- Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling
- Inverted Detection and Control in Steering Vectors
- When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars
- Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
- Necessary, Decodable and Reversible, Yet Not Transferable: A Stress Test for Attention-Head Role Claims
- When More Becomes Less: Position-Dependent Repetition Effects in Language Models
- Do LLMs Know What Is Private Internally? Probing and Steering Contextual Privacy Norms in Large Language Model Representations
- Evaluating Vision Language Models (VLMs) for Radiology: A Comprehensive Analysis
- Decoding Vision Transformers: the Diffusion Steering Lens
- Detecting Safety Training Modification in Language Models via Activation Analysis
- Support Operation Factorization: Compositional Readout of Frozen Vision Encoders under Controlled Interventions
- Visual Grounding in Zero-Shot Vision-Language Control
- The Ignition Index: Measuring Global Workspace Dynamics in Language Models
- Probing then Editing Response Personality of Large Language Models
- Linguistic Interpretability of Transformer-based Language Models: a systematic review
- On the Effectiveness and Generalization of Race Representations for Debiasing High-Stakes Decisions
Related