Shape Happens: Automatic Feature Manifold Discovery in LLMs via Supervised Multi-Dimensional Scaling
2025/10/01 by Tiblias, Federico, Bigoulaeva, Irina, Niu, Jingcheng +2
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2510.01025
Abstract
The linear representation hypothesis states that language models (LMs) encode concepts as directions in their latent space, forming organized, multidimensional manifolds. Prior efforts focus on discovering specific geometries for specific features, and thus lack generalization. We introduce Supervised Multi-Dimensional Scaling (SMDS), a model-agnostic method to automatically discover feature manifolds. We apply SMDS to temporal reasoning as a case study, finding that different features form various geometric structures such as circles, lines, and clusters. SMDS reveals many insights on these structures: they consistently reflect the properties of the concepts they represent; are stable across model families and sizes; actively support reasoning in models; and dynamically reshape in response to context changes. Together, our findings shed light on the functional role of feature manifolds, supporting a model of entity-based reasoning in which LMs encode and transform structured representations.
Citations
- The Origins of Representation Manifolds in Large Language Models
- Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs
- TRAVELER: A Benchmark for Evaluating Temporal Reasoning across Vague, Implicit and Explicit References
- The Reasoning-Memorization Interplay in Language Models Is Mediated by a Single Direction
- RTD-Lite: Scalable Topological Analysis for Comparing Weighted Graphs in Learning Tasks
- Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
- What is a Number, That a Large Language Model May Know It?
- AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
- Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
- Applying sparse autoencoders to unlearn knowledge in language models
- The Geometry of Numerical Reasoning: Language Models Compare Numeric Properties in Linear Subspaces
- Language Models Encode Numbers Using Digit Representations in Base 10
- Concept Space Alignment in Multilingual LLMs
- A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
- Representational Analysis of Binding in Language Models
- Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
- The Geometry of Categorical and Hierarchical Concepts in Large Language Models
- Not All Language Model Features Are One-Dimensionally Linear
- A Closer Look at the Limitations of Instruction Tuning
- The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning
- TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models
- The Linear Representation Hypothesis and the Geometry of Large Language Models
- How do Language Models Bind Entities in Context?
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- Language Models Represent Space and Time
- From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Bias and Fairness in Large Language Models: A Survey
- Bias and Fairness in Large Language Models: A Survey
- Why do universal adversarial attacks work on large language models?: Geometry might be the answer
- Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language Models
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
- LIMA: Less Is More for Alignment
- Finding Neurons in a Haystack: Case Studies with Sparse Probing
- Towards Automated Circuit Discovery for Mechanistic Interpretability
- Zero-shot Temporal Relation Extraction with ChatGPT
- Large Language Models Can Be Easily Distracted by Irrelevant Context
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
- Scaling Instruction-Finetuned Language Models
- The Geometry of Multilingual Language Model Representations
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Finetuned Language Models Are Zero-Shot Learners
- A Dataset for Answering Time-Sensitive Questions
- Probing Classifiers: Promises, Shortcomings, and Advances
- Probing Classifiers: Promises, Shortcomings, and Advances
- Multidimensional Scaling, Sammon Mapping, and Isomap: Tutorial and Survey
- Emerging Cross-lingual Structure in Pretrained Language Models
- Designing and Interpreting Probes with Control Tasks
- "Going on a vacation" takes longer than "Going for a walk": A Study of Temporal Commonsense Understanding
- BERTScore: Evaluating Text Generation with BERT
- Understanding intermediate layers using linear classifier probes
- Aligning Large Language Models with Human: A Survey
Related