Discovering Latent Knowledge in Language Models Without Supervision
2022/12/07 by Collin Burns, Haotian Ye, Burns, Collin +5 · 5 voices · 222 citations
Computer Science · #Artificial intelligence #Computer science #Consistency (knowledge bases) #Ground truth #Language model #Linguistics #Machine learning #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Natural language processing #Negation #Programming language #Question answering #Space (punctuation) #Statement (logic) #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2212.03827
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2022/12/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Abstract
Existing techniques for training language models can be misaligned with the truth: if we train models with imitation learning, they may reproduce errors that humans make; if we train them to generate text that humans rate highly, they may output errors that human evaluators can't detect. We propose circumventing this issue by directly finding latent knowledge inside the internal activations of a language model in a purely unsupervised way. Specifically, we introduce a method for accurately answering yes-no questions given only unlabeled model activations. It works by finding a direction in activation space that satisfies logical consistency properties, such as that a statement and its negation have opposite truth values. We show that despite using no supervision and no model outputs, our method can recover diverse knowledge represented in large language models: across 6 models and 10 question-answering datasets, it outperforms zero-shot accuracy by 4% on average. We also find that it cuts prompt sensitivity in half and continues to maintain high accuracy even when models are prompted to generate incorrect answers. Our results provide an initial step toward discovering what language models know, distinct from what they say, even when we don't have access to explicit ground truth labels.
Cited by
- Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study
- Dissociating the Internal Representations of Sycophancy in LLMs
- Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models
- Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making
- Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model
- Reading Calibrated Uncertainty from Language Model Trajectories
- The Anatomy of a Truth Direction: Knowledge-Dependent Dimensionality, a Relational Law, and a Convergent Category Geometry in Small Language Models
- TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics
- Gotta Catch them all: the modes of Sycophancy
- From Sycophancy to Deception: A Unified Taxonomy for LLM Spontaneous Misalignment
- How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?
- Belief-reality separation lives in routing over a shared value slot in language models
- Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes
- Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins
- Diversity-Oriented Fine-Tuning for Uncertainty-Based Hallucination Detection
- Diagnosing Correctness Probes under Self-Judgement Confounding
- Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models
- Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors
- When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space
- Linear representations of grammaticality in neural language models
- Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade
- Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models
- Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs
- Securing Multimodal AI through Internal Information Decomposition
- Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models
- Latent Introspection: Models Can Detect Prior Concept Injections
- Linear representations in language models can change dramatically over a conversation
- Unsupervised Elicitation of Language Models
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
- ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions
- Towards Understanding Steering Strength
- Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
- TRE: Training-Free Hallucination Detection for Diffusion Language Models
- Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs
- Do Models Fake Alignment Without Clear Consequences?
- Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits
- The Deleuzian Representation Hypothesis
- LYNX: Learning Dynamic Exits for Confidence-Controlled Reasoning
- HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs
- SPAD: Seven-Source Token Probability Attribution with Syntactic Aggregation for Detecting Hallucinations in RAG
- ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior
- Group Selection as a Safeguard Against AI Substitution
- Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
- Query-Level Uncertainty in Large Language Models
- Are LLMs Good Safety Agents or a Propaganda Engine?
- Difficulties with Evaluating a Deception Detector for AIs
- REFLEX: Self-Refining Explainable Fact-Checking via Disentangling Truth into Style and Substance
- Unboxing the Black Box: Mechanistic Interpretability for Algorithmic Understanding of Neural Networks
- Polarity-Aware Probing for Quantifying Latent Alignment in Language Models
- Beyond Tokens in Language Models: Interpreting Activations through Text Genre Chunks
- Rethinking Deep Alignment Through The Lens Of Incomplete Learning
- Spilling the Beans: Teaching LLMs to Self-Report Their Hidden Objectives
- Self-Evaluating LLMs for Multi-Step Tasks: Stepwise Confidence Estimation for Failure Detection
- You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
- Beyond Correctness: Confidence-Aware Reward Modeling for Enhancing Large Language Model Reasoning
- Interpreting Multi-Attribute Confounding through Numerical Attributes in Large Language Models
- How LLMs Detect and Correct Their Own Errors: The Role of Internal Confidence Signals
- LLM Probing with Contrastive Eigenproblems: Improving Understanding and Applicability of CCS
- ThoughtProbe: Classifier-Guided LLM Thought Space Exploration via Probing Representations
- ParaScopes: What do Language Models Activations Encode About Future Text?
- Layer of Truth: Probing Belief Shifts under Continual Pre-Training Poisoning
- The Refusal Residue: When Probes Catch Alignment Faking and When They Don't
- How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs
- Subliminal Corruption: Mechanisms, Thresholds, and Interpretability
- Emergence of Linear Truth Encodings in Language Models
- Beyond "Hallucinations": A Framework for Stable Human-AI Reasoning
- AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefs
- Bolster Hallucination Detection via Prompt-Guided Data Augmentation
- A Two-Step, Multidimensional Account of Deception in Language Models
- LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
- A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis
- How do LLMs Compute Verbal Confidence
- SIMBA UQ: Similarity-Based Aggregation for Uncertainty Quantification in Large Language Models
- SAFE: Multitask Failure Detection for Vision-Language-Action Models
- Neologism Learning for Controllability and Self-Verbalization
- On the Convergence of Moral Self-Correction in Large Language Models
- The Geometry of Truth: Layer-wise Semantic Dynamics for Hallucination Detection in Large Language Models
- LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
- Beyond Token Probes: Hallucination Detection via Activation Tensors with ACT-ViT
- Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
- TraceDet: Hallucination Detection from the Decoding Trace of Diffusion Large Language Models
- Interpret, prune and distill Donut : towards lightweight VLMs for VQA on document
- Neural Message-Passing on Attention Graphs for Hallucination Detection
- Toward Preference-aligned Large Language Models via Residual-based Model Steering
- SafeSteer: Adaptive Subspace Steering for Efficient Jailbreak Defense in Vision-Language Models
- From Input Perception to Predictive Insight: Modeling Model Blind Spots Before They Become Errors
- The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models
- Training-free Truthfulness Detection via Sparse MLP Value Vectors
- nDNA -- the Semantic Helix of Artificial Cognition
- Representation-based Broad Hallucination Detectors Fail to Generalize Out of Distribution
- SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
- RepIt: Steering Language Models with Concept-Specific Refusal Vectors
- HalluField: Detecting LLM Hallucinations via Field-Theoretic Modeling
- No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
- Unsupervised Hallucination Detection by Inspecting Reasoning Processes
- AI Wellbeing
- Cross-Layer Attention Probing for Fine-Grained Hallucination Detection
- Can LLMs Lie? Investigation beyond Hallucination
- Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions
- Enhancing Uncertainty Estimation in LLMs with Expectation of Aggregated Internal Belief
- Two Causes, Not One: Rethinking Omission and Fabrication Hallucinations in MLLMs
- STARE at the Structure: Steering ICL Exemplar Selection with Structural Alignment
- Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
- GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity
- Continuously Steering LLMs Sensitivity to Contextual Knowledge with Proxy Models
- Caught in the Act: a mechanistic approach to detecting deception
- Quantized but Deceptive? A Multi-Dimensional Truthfulness Evaluation of Quantized LLMs
- Universal and Transferable Adversarial Attack on Large Language Models Using Exponentiated Gradient Descent
- Magical: Medical Lay Language Generation via Semantic Invariance and Layperson-tailored Adaptation
- Never compromise with vulnerabilities: a comprehensive survey on AI governance
- LLM Unlearning using Gradient Ratio-Based Influence Estimation and Noise Injection
- Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
- Let's Measure Information Step-by-Step: LLM-Based Evaluation Beyond Vibes
- AI-AI Bias: large language models favor communications generated by large language models
- ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges
- Balancing Stylization and Truth via Disentangled Representation Steering
- A Survey on Data Security in Large Language Models
- Depth Gives a False Sense of Privacy: LLM Internal States Inversion
- How Does Controllability Emerge In Language Models During Pretraining?
- Foundations of Interpretable Models
- Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
- LENS: Learning Ensemble Confidence from Neural States for Multi-LLM Answer Integration
- A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
- NPO: Learning Alignment and Meta-Alignment through Structured Human Feedback
- Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework
- Can We Predict Alignment Before Models Finish Thinking? Towards Monitoring Misaligned Reasoning Models
- Internal Value Alignment in Large Language Models through Controlled Value Vector Activation
- Deep Hidden Cognition Facilitates Reliable Chain-of-Thought Reasoning
- The Geometries of Truth Are Orthogonal Across Tasks
- The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
- Beyond the Linear Separability Ceiling: Aligning Representations in VLMs
- May I have your Attention? Breaking Fine-Tuning based Prompt Injection Defenses using Architecture-Aware Attacks
- CoRE: Enhancing Metacognition with Label-free Self-evaluation in LRMs
- Data Supplement to the paper "Intervening to Learn and Compose Causally Disentangled Representations"
- On the Semantics of Large Language Models
- PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization
- Probing and Steering Evaluation Awareness of Language Models
- Feature Integration Spaces: Joint Training Reveals Dual Encoding in Neural Network Representations
- TuCo: Measuring the Contribution of Fine-Tuning to Individual Responses of LLMs
- KScope: A Framework for Characterizing the Knowledge Status of Language Models
- Beyond Autocomplete: Designing CopilotLens Towards Transparent and Explainable AI Coding Agents
- Mechanistic Interpretability Needs Philosophy
- No Training Wheels: Steering Vectors for Bias Correction at Inference Time
- Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack)
- The Role of Model Confidence on Bias Effects in Measured Uncertainties for Vision-Language Models
- From Emergence to Control: Probing and Modulating Self-Reflection in Language Models
- Convergent Linear Representations of Emergent Misalignment
- Detecting High-Stakes Interactions with Activation Probes
- Bridging External and Parametric Knowledge: Mitigating Hallucination of LLMs with Shared-Private Semantic Synergy in Dual-Stream Knowledge
- Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models
- Cracking the Code: Enhancing Implicit Hate Speech Detection through Coding Classification
- When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
- Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
- One SPACE to Rule Them All: Jointly Mitigating Factuality and Faithfulness Hallucinations in LLMs
- AI Agent Behavioral Science
- Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations
- Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning
- Expanding before Inferring: Enhancing Factuality in Large Language Models through Premature Layers Interpolation
- Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes
- Probing the Geometry of Truth: Consistency and Generalization of Truth Directions in LLMs Across Logical Transformations and Question Answering Tasks
- SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs
- Probing Neural Topology of Large Language Models
- Revisiting Epistemic Markers in Confidence Estimation: Can Markers Accurately Reflect Large Language Models' Uncertainty?
- COSMIC: Generalized Refusal Direction Identification in LLM Activations
- MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs
- HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMs
- Revisiting Uncertainty Estimation and Calibration of Large Language Models
- Active Layer-Contrastive Decoding Reduces Hallucination in Large Language Model Generation
- Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling
- Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders
- JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language Models
- Response Uncertainty and Probe Modeling: Two Sides of the Same Coin in LLM Interpretability?
- But what is your honest answer? Aiding LLM-judges with honest alternatives using steering vectors
- Activation Control for Efficiently Eliciting Long Chain-of-thought Ability of Language Models
- Too Consistent to Detect: A Study of Self-Consistent Errors in LLMs
- LLaMAs Have Feelings Too: Unveiling Sentiment and Emotion Representations in LLaMA Models Through Probing
- Transformer Copilot: Learning from The Mistake Log in LLM Fine-tuning
- When Do LLMs Admit Their Mistakes? Understanding The Role Of Model Belief In Retraction
- Void in Language Models
- Mining Intrinsic Rewards from LLM Hidden States for Efficient Best-of-N Sampling
- Truth Neurons
- Exploring the generalization of LLM truth directions on conversational formats
- Can LLMs Reliably Self-Report Adversarial Prefills, and How?
- Multi-agents based User Values Mining for Recommendation
- Embracing Anisotropy: Turning Massive Activations into Interpretable Control Knobs for Large Language Models
- Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure
- Readable but Not Controllable: Neuron-Level Evidence for Medical LLM Hallucination
- Me, Myself, and π : Evaluating and Explaining LLM Introspection
- The Value Axis: Language Models Encode Whether They're on the Right Track
- When Role-playing, Do Models Believe What They Say?
- Prompt Injection as Role Confusion
- Positive Alignment: Artificial Intelligence for Human Flourishing
- LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations
- Manifold-Guided Attention Steering
- Weak-to-Strong Generalization via Direct On-Policy Distillation
- Subliminal Learning Is Steering Vector Distillation
- Can LLMs Introspect? A Reality Check
- Monitoring Emergent Reward Hacking During Generation via Internal Activations
- The Truthfulness Spectrum Hypothesis
- Towards Long Context Hallucination Detection
- Surgical Repair of Insecure Code Generation in LLMs
- ProbeLogits: Kernel-Level LLM Inference Primitives for AI-Native Operating Systems
- The Spectral Geometry of Thought: Phase Transitions, Instruction Reversal, Token-Level Dynamics, and Perfect Correctness Prediction in How Transformers Reason
- Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks
- Language Models Encode the Contextual Truth of Propositions
- Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition
- Test, then Route: How Language Models Execute In-Context Conditional Rules Across Models and Languages
- Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
- Among Us: A Sandbox for Measuring and Detecting Agentic Deception
- Necessary, Decodable and Reversible, Yet Not Transferable: A Stress Test for Attention-Head Role Claims
- Exploring How LLMs Capture and Represent Domain-Specific Knowledge
- Object-Level Verbalized Confidence Calibration in Vision-Language Models via Semantic Perturbation
- Functional Abstraction of Knowledge Recall in Large Language Models
- The Geometry of Self-Verification in a Task-Specific Reasoning Model
- Hypothetical Documents or Knowledge Leakage? Rethinking LLM-based Query Expansion
- Position: It's Time to Optimize LLMs for Self-Consistency
- Enhancing Mathematical Reasoning in Large Language Models with Self-Consistency-Based Hallucination Detection
- HalluShift: Measuring Distribution Shifts towards Hallucination Detection in LLMs
- Alleviating the Fear of Losing Alignment in LLM Fine-tuning
- Robust Hallucination Detection in LLMs via Adaptive Token Selection
- ThoughtProbe: Classifier-Guided Thought Space Exploration Leveraging LLM Intrinsic Reasoning
- Paul Christiano [wikipedia]
Discussions
- Discovering latent knowledge in language models without supervision [hn, 149 points, 85 comments]
- syntax rules alone are not sufficient to perform the tasks LLMs perform here is a paper on knowledge representation inside LLMs: arxiv.org/abs/2212.03827 [bsky, 3 points, 0 comments]
- Discovering Latent Knowledge in Language Models Without Supervision [lobsters, 3 points, 1 comments]
- There’s been research since 2022 that suggests LLMs have internal representations re whether certain statements are true. It’s an entire research field in interpretability studies arxiv.org/abs/2212.0 [bsky, 2 points, 0 comments]
- Discovering Latent Knowledge in Language Models Without Supervision [hn, 2 points, 0 comments]
Related