External Invariants: A Cryptographic Trust Architecture for Institutional AI Inference
2024/06/17 by Andy Arditi, Arditi, Andy, Oscar Obeso +11 · 20 voices · 192 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2406.11717
openalex publication_date 2024/06/17 · openalex created_date 2024/06/19 · openalex updated_date 2026/07/28
Abstract
Conversational large language models are fine-tuned for both instruction-following and safety, resulting in models that obey benign requests but refuse harmful ones. While this refusal behavior is widespread across chat models, its underlying mechanisms remain poorly understood. In this work, we show that refusal is mediated by a one-dimensional subspace, across 13 popular open-source chat models up to 72B parameters in size. Specifically, for each model, we find a single direction such that erasing this direction from the model's residual stream activations prevents it from refusing harmful instructions, while adding this direction elicits refusal on even harmless instructions. Leveraging this insight, we propose a novel white-box jailbreak method that surgically disables refusal with minimal effect on other capabilities. Finally, we mechanistically analyze how adversarial suffixes suppress propagation of the refusal-mediating direction. Our findings underscore the brittleness of current safety fine-tuning methods. More broadly, our work showcases how an understanding of model internals can be leveraged to develop practical methods for controlling model behavior.
Cited by
- V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure
- Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study
- Context Is King: How In-Context Specification Shapes the Geometry of Concepts
- Defense Against LLM Backdoors using Critical Neuron Isolation Pruning
- Sound Probabilistic Safety Bounds for Large Language Models
- The Ethics of Autonomous AI Agents for Offensive Security
- Geometric Configurations of Perturbed Jailbreak Prompts
- Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs
- CircuitKIT : Circuit Discovery, Evaluation, and Application Toolkit for Mechanistic Interpretability
- Matching Ranks Over Probability Yields Truly Deep Safety Alignment
- Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control
- Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak
- Estimating Rare Events in Language Models with Proper Evaluation
- Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary
- Probing the Difficulty Perception Mechanism of Large Language Models
- When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space
- Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control
- ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs
- DiMaS: Distribution Matching for Steering Vision-Language-Action Models
- Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior
- TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment
- Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models
- Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal
- DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions
- Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models
- Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate
- Mechanisms of Introspective Awareness
- LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
- Interpreting Physics in Video World Models
- Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework
- Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
- Base Models Know How to Reason, Thinking Models Learn When
- Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs
- Steering Large Language Models for Machine Translation Personalization
- Dialz: A Python Toolkit for Steering Vectors
- Training large language models on narrow tasks can lead to broad misalignment
- Rigorous Interpretation Is a Form of Evaluation
- Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
- Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
- ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions
- Towards Understanding Steering Strength
- Do LLMs Know Their Vulnerable Scenarios?
- Where Steering Signals Come From: Activation Source Selection in Activation Steering
- Many-body Tipping Dynamics of ChatGPT-like AIs
- Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT
- Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering
- Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks
- How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift
- xMIx: High-Performance Serving-Time Platform for Mechanistic Interpretability Apps
- Reference Feature Atlases for Mechanistic Auditing of Language Models
- Gabliteration: Adaptive Multi-Directional Neural Weight Modification for Selective Behavioral Alteration in Large Language Models
- "Even GPT Can Reject Me": Conceptualizing Abrupt Refusal Secondary Harm (ARSH) and Reimagining Psychological AI Safety with Compassionate Completion Standard (CCS)
- Linear Personality Probing and Steering in LLMs: A Big Five Study
- Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
- From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
- Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
- Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation
- Sparse Concept Anchoring for Interpretable and Controllable Neural Representations
- Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
- MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents
- Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
- SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation
- Graph-Regularized Sparse Autoencoders for LLM Safety Steering
- ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models
- In-Context Representation Hijacking
- Label Forensics: Interpreting Hard Labels in Black-Box Text Classifier
- When Safety Blocks Sense: Measuring Semantic Confusion in LLM Refusals
- Are LLMs Good Safety Agents or a Propaganda Engine?
- Decomposed Trust: Exploring Privacy, Adversarial Robustness, Fairness, and Ethics of Low-Rank LLMs
- Representation Interventions Enable Lifelong Unstructured Knowledge Control
- SAGE: An Agentic Explainer Framework for Interpreting SAE Features in Language Models
- Physics Steering: Causal Control of Cross-Domain Concepts in a Physics Foundation Model
- TRANSPORTER: Transferring Visual Semantics from VLM Manifolds
- Curvature-Aware Safety Restoration In LLMs Fine-Tuning
- Steering in the Shadows: Causal Amplification for Activation Space Attacks in Large Language Models
- Unified Defense for Large Language Models against Jailbreak and Fine-Tuning Attacks in Education
- N-GLARE: An Non-Generative Latent Representation-Efficient LLM Safety Evaluator
- Synthetic Eggs in Many Baskets: The Impact of Synthetic Data Diversity on LLM Fine-Tuning
- Rethinking Deep Alignment Through The Lens Of Incomplete Learning
- SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models
- Why does weak-OOD help? A Further Step Towards Understanding Jailbreaking VLMs
- Investigating CoT Monitorability in Large Reasoning Models
- Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment
- You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
- Steering Language Models with Weight Arithmetic
- Model Merging Improves Zero-Shot Generalization in Bioacoustic Foundation Models
- Minimal and Mechanistic Conditions for Behavioral Self-Awareness in LLMs
- Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
- AdversariaLLM: A Unified and Modular Toolbox for LLM Robustness Research
- Silenced Biases: The Dark Side LLMs Learned to Refuse
- ShadowLogic: Backdoors in Any Whitebox LLM
- Red-teaming Activation Probes using Prompted LLMs
- Neural Transparency: Mechanistic Interpretability Interfaces for Anticipating Model Behaviors for Personalized AI
- RepV: Safety-Separable Latent Spaces for Scalable Neurosymbolic Plan Verification
- Chain-of-Thought Hijacking
- Angular Steering: Behavior Control via Rotation in Activation Space
- Adaptively Robust LLM Monitoring via Activation Watermarking
- Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
- Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis
- ToxScreen: Detecting Whether an LLM Has Been Poisoned
- State-Dependent Safety Failures in Multi-Turn Language Model Interaction
- Navigating through the hidden embedding space: steering LLMs to improve mental health assessment
- Sequences of Logits Reveal the Low Rank Structure of Language Models
- Can Aha Moments Be Fake? Identifying True and Decorative Thinking Steps in Chain-of-Thought
- Do Stop Me Now: Detecting Boilerplate Responses with a Single Iteration
- Mapping Faithful Reasoning in Language Models
- Modeling Hierarchical Thinking in Large Reasoning Models
- Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
- Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training
- Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
- Agentic Reinforcement Learning for Search Misaligns Instruction-Tuning
- SARSteer: Safeguarding Large Audio-Language Models via Safe-Ablated Refusal Steering
- Who's Asking? Simulating Role-Based Questions for Conversational AI Evaluation
- To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models
- In-Distribution Steering: Balancing Control and Coherence in Language Model Generation
- Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment
- The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs
- RAID: Refusal-Aware and Integrated Decoding for Jailbreaking LLMs
- Keep Calm and Avoid Harmful Content: Concept Alignment and Latent Manipulation Towards Safer Answers
- Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
- Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research
- The Unintended Trade-off of AI Alignment:Balancing Hallucination Mitigation and Safety in LLMs
- PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector Algebra
- How Language Models Conflate Logical Validity with Plausibility: A Representational Analysis of Content Effects
- LatentBreak: Jailbreaking Large Language Models through Latent Space Feedback
- Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?
- Prototype-Based Dynamic Steering for Large Language Models
- Read the Scene, Not the Script: Outcome-Aware Safety for LLMs
- Activation Steering with a Feedback Controller
- MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
- A Granular Study of Safety Pretraining under Model Abliteration
- Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
- AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
- ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
- Defeating Cerberus: Concept-Guided Privacy-Leakage Mitigation in Multimodal Language Models
- EasySteer: A Unified Framework for High-Performance and Extensible LLM Steering
- DiffuGuard: How Intrinsic Safety is Lost and Found in Diffusion Large Language Models
- Toward Preference-aligned Large Language Models via Residual-based Model Steering
- Beyond the Exploration-Exploitation Trade-off: A Hidden State Approach for LLM Reasoning in RLVR
- Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
- Enhancing LLM Steering through Sparse Autoencoder-Based Vector Refinement
- Train Once, Answer All: Many Pretraining Experiments for the Cost of One
- Knowledge distillation through geometry-aware representational alignment
- IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
- Detecting (Un)answerability in Large Language Models with Linear Directions
- The Rogue Scalpel: Activation Steering Compromises LLM Safety
- Active Attacks: Red-teaming LLMs via Adaptive Environments
- Hallucination reduction with CASAL: Contrastive Activation Steering For Amortized Learning
- Silent Tokens, Loud Effects: Padding in LLMs
- Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters
- Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game
- Aligning Recommendations with User Popularity Preferences
- Steering When Necessary: Flexible Steering Large Language Models with Backtracking
- DISCO: Disentangled Communication Steering for Large Language Models
- SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection
- Representation-based Broad Hallucination Detectors Fail to Generalize Out of Distribution
- Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction
- LLM Jailbreak Detection for (Almost) Free!
- AdaptiveK Sparse Autoencoders: Dynamic Sparsity Allocation for Interpretable LLM Representations
- Programmable Cognitive Bias in Social Agents
- SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
- RepIt: Steering Language Models with Concept-Specific Refusal Vectors
- LLM Assertiveness can be Mechanistically Decomposed into Emotional and Logical Components
- Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks
- Unveiling the Latent Directions of Reflection in Large Language Models
- So let's replace this phrase with insult... Lessons learned from generation of toxic texts with LLMs
- Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint
- Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift
- Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
- MEUV: Achieving Fine-Grained Capability Activation in Large Language Models via Mutually Exclusive Unlock Vectors
- Unlearning That Lasts: Utility-Preserving, Robust, and Almost Irreversible Forgetting in LLMs
- From Indirect Object Identification to Syllogisms: Exploring Binary Mechanisms in Transformer Circuits
- ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
- Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
- Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
- Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
- Safety Alignment Should Be Made More Than Just A Few Attention Heads
- Sycophancy as compositions of Atomic Psychometric Traits
- Learning to Steer: Input-dependent Steering for Multimodal LLMs
- Consiglieres in the Shadow: Understanding the Use of Uncensored Large Language Models in Cybercrimes
- Mitigating Jailbreaks with Intent-Aware LLMs
- Enhancing Supervised Composed Image Retrieval via Reasoning-Augmented Representation Engineering
- SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering
- Human-Alignment and Calibration of Inference-Time Uncertainty in Large Language Models
- VISOR: Visual Input-based Steering for Output Redirection in Vision-Language Models
- Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
- Semantic Structure in Large Language Model Embeddings
- Don't Overthink It: A Survey of Efficient R1-style Large Reasoning Models
- Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders
- Multilingual Political Views of Large Language Models: Identification and Steering
- The Blessing and Curse of Dimensionality in Safety Alignment
- Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment
Discussions
- Refusal in language models is mediated by a single direction [hn, 209 points, 44 comments]
- Refusal in Language Models Is Mediated by a Single Direction [hn, 118 points, 45 comments]
- For example, Arditi et al. (arxiv.org/abs/2406.11717) argued that refusal is mediated by a single "direction," or linear subspace, in many language models. But when we applied their method on a RWKV m [bsky, 4 points, 1 comments]
- My take: models are capable of learning some patterns, and not others. But, without a causal experiment, we don't know a priori know what these patterns are. This is my favorite recent example that e [bsky, 2 points, 0 comments]
- Training LLMs includes teaching them to sometimes respond "I'm sorry, but I can't answer that". AI research calls this "refusal" and it is one of many separable proto-concepts in these systems. This A [bsky, 2 points, 1 comments]
- Yes, the technique they used is the same technique that heretic uses! (From https://arxiv.org/abs/2406.11717) [bsky, 1 points, 0 comments]
- Refusal in language models is mediated by a single direction (arxiv.org) Main Link | Discussion [bsky, 1 points, 0 comments]
- Refusal in Language Models Is Mediated by a Single Direction https://arxiv.org/abs/2406.11717 [comments] [70 points] [bsky, 0 points, 0 comments]
- I *suspect* you run into an issue similar to the one described in part 5 of "Refusal in LMs is mediated by a single direction" arxiv.org/abs/2406.11717 where "refuse injected commands" becomes part of [bsky, 0 points, 1 comments]
- Re: Jailbreaks I suspect mechanistic analysis of various jailbreaks will give us a better idea of how to prevent them. I feel like there ultimately isn't a fix for something like this prefix attack th [bsky, 0 points, 1 comments]
- Models refusing requests through one linear direction in activation space? That's surprisingly simple – and probably why jailbreaks work so well. https://arxiv.org/abs/2406.11717 [bsky, 0 points, 1 comments]
- https://arxiv.org/abs/2406.11717 大規模言語モデルの拒否挙動は単一方向で制御可能であることを発見しました。 この方向を削除すると拒否がなくなり、追加すると無害な指示も拒否します。 安全対策の脆弱性を示唆し、モデル内部理解の重要性を強調しています。 [bsky, 0 points, 0 comments]
- Refusal in Language Models Is Mediated by a Single Direction https:// arxiv.org/abs/2406.11717 # arxiv [mastodon, 0 points, 0 comments]
- Refusal in Language Models Is Mediated by a Single Direction https://arxiv.org/abs/2406.11717 [bsky, 0 points, 0 comments]
- Refusal in Language Models Is Mediated by a Single Direction https://arxiv.org/abs/2406.11717 (https://news.ycombinator.com/item?id=47986136) [bsky, 0 points, 0 comments]
- Refusal in Language Models Is Mediated by a Single Direction https://arxiv.org/abs/2406.11717 (https://news.ycombinator.com/item?id=47986136) [bsky, 0 points, 0 comments]
- Refusal in Language Models Is Mediated by a Single Direction [bsky, 0 points, 0 comments]
- Refusal in Language Models Is Mediated by a Single Direction #HackerNews https://arxiv.org/abs/2406.11717 [bsky, 0 points, 0 comments]
- Refusal in Language Models Is Mediated by a Single Direction https://arxiv.org/abs/2406.11717 https://news.ycombinator.com/item?id=47986136 [bsky, 0 points, 0 comments]
- Refusal in Language Models Is Mediated by a Single Direction arxiv.org/pdf/2406.11717 [bsky, 0 points, 0 comments]
Related