External Invariants: A Cryptographic Trust Architecture for Institutional AI Inference
2024/06/17 by Andy Arditi, Arditi, Andy, Oscar Obeso +11 · 20 voices · 362 citations
Computer Science · Psychology · #Computer science #Linguistics #Natural Language Processing Techniques #Philosophy #Psychology #Topic Modeling #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2406.11717
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/06/17 · openalex created_date 2024/06/19 · openalex updated_date 2026/07/28
Abstract
Conversational large language models are fine-tuned for both instruction-following and safety, resulting in models that obey benign requests but refuse harmful ones. While this refusal behavior is widespread across chat models, its underlying mechanisms remain poorly understood. In this work, we show that refusal is mediated by a one-dimensional subspace, across 13 popular open-source chat models up to 72B parameters in size. Specifically, for each model, we find a single direction such that erasing this direction from the model's residual stream activations prevents it from refusing harmful instructions, while adding this direction elicits refusal on even harmless instructions. Leveraging this insight, we propose a novel white-box jailbreak method that surgically disables refusal with minimal effect on other capabilities. Finally, we mechanistically analyze how adversarial suffixes suppress propagation of the refusal-mediating direction. Our findings underscore the brittleness of current safety fine-tuning methods. More broadly, our work showcases how an understanding of model internals can be leveraged to develop practical methods for controlling model behavior.
Cited by
- V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure
- Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study
- Context Is King: How In-Context Specification Shapes the Geometry of Concepts
- Defense Against LLM Backdoors using Critical Neuron Isolation Pruning
- Sound Probabilistic Safety Bounds for Large Language Models
- The Ethics of Autonomous AI Agents for Offensive Security
- Geometric Configurations of Perturbed Jailbreak Prompts
- Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs
- CircuitKIT : Circuit Discovery, Evaluation, and Application Toolkit for Mechanistic Interpretability
- Matching Ranks Over Probability Yields Truly Deep Safety Alignment
- Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control
- Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak
- Estimating Rare Events in Language Models with Proper Evaluation
- Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary
- Probing the Difficulty Perception Mechanism of Large Language Models
- When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space
- Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control
- ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs
- DiMaS: Distribution Matching for Steering Vision-Language-Action Models
- Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior
- TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment
- Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models
- Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal
- DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions
- Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models
- Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate
- Mechanisms of Introspective Awareness
- LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
- Interpreting Physics in Video World Models
- Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework
- Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
- Base Models Know How to Reason, Thinking Models Learn When
- Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs
- Steering Large Language Models for Machine Translation Personalization
- Dialz: A Python Toolkit for Steering Vectors
- Training large language models on narrow tasks can lead to broad misalignment
- Rigorous Interpretation Is a Form of Evaluation
- Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
- Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
- ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions
- Towards Understanding Steering Strength
- Do LLMs Know Their Vulnerable Scenarios?
- Where Steering Signals Come From: Activation Source Selection in Activation Steering
- Many-body Tipping Dynamics of ChatGPT-like AIs
- Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT
- Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering
- Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks
- How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift
- xMIx: High-Performance Serving-Time Platform for Mechanistic Interpretability Apps
- Reference Feature Atlases for Mechanistic Auditing of Language Models
- Gabliteration: Adaptive Multi-Directional Neural Weight Modification for Selective Behavioral Alteration in Large Language Models
- "Even GPT Can Reject Me": Conceptualizing Abrupt Refusal Secondary Harm (ARSH) and Reimagining Psychological AI Safety with Compassionate Completion Standard (CCS)
- Linear Personality Probing and Steering in LLMs: A Big Five Study
- Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
- From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
- Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
- Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation
- Sparse Concept Anchoring for Interpretable and Controllable Neural Representations
- Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
- MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents
- Unforgotten Safety: Preserving Safety Alignment of Large Language Models with Continual Learning
- SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation
- Graph-Regularized Sparse Autoencoders for LLM Safety Steering
- ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models
- In-Context Representation Hijacking
- Label Forensics: Interpreting Hard Labels in Black-Box Text Classifier
- When Safety Blocks Sense: Measuring Semantic Confusion in LLM Refusals
- Are LLMs Good Safety Agents or a Propaganda Engine?
- Decomposed Trust: Exploring Privacy, Adversarial Robustness, Fairness, and Ethics of Low-Rank LLMs
- Representation Interventions Enable Lifelong Unstructured Knowledge Control
- SAGE: An Agentic Explainer Framework for Interpreting SAE Features in Language Models
- Physics Steering: Causal Control of Cross-Domain Concepts in a Physics Foundation Model
- DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt
- TRANSPORTER: Transferring Visual Semantics from VLM Manifolds
- Curvature-Aware Safety Restoration In LLMs Fine-Tuning
- Steering in the Shadows: Causal Amplification for Activation Space Attacks in Large Language Models
- Unified Defense for Large Language Models against Jailbreak and Fine-Tuning Attacks in Education
- N-GLARE: An Non-Generative Latent Representation-Efficient LLM Safety Evaluator
- Synthetic Eggs in Many Baskets: The Impact of Synthetic Data Diversity on LLM Fine-Tuning
- Rethinking Deep Alignment Through The Lens Of Incomplete Learning
- SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models
- Why does weak-OOD help? A Further Step Towards Understanding Jailbreaking VLMs
- Investigating CoT Monitorability in Large Reasoning Models
- Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment
- You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
- Steering Language Models with Weight Arithmetic
- Model Merging Improves Zero-Shot Generalization in Bioacoustic Foundation Models
- Minimal and Mechanistic Conditions for Behavioral Self-Awareness in LLMs
- Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
- AdversariaLLM: A Unified and Modular Toolbox for LLM Robustness Research
- Silenced Biases: The Dark Side LLMs Learned to Refuse
- ShadowLogic: Backdoors in Any Whitebox LLM
- Red-teaming Activation Probes using Prompted LLMs
- Neural Transparency: Mechanistic Interpretability Interfaces for Anticipating Model Behaviors for Personalized AI
- RepV: Safety-Separable Latent Spaces for Scalable Neurosymbolic Plan Verification
- Chain-of-Thought Hijacking
- Angular Steering: Behavior Control via Rotation in Activation Space
- Adaptively Robust LLM Monitoring via Activation Watermarking
- Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
- Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis
- ToxScreen: Detecting Whether an LLM Has Been Poisoned
- State-Dependent Safety Failures in Multi-Turn Language Model Interaction
- Navigating through the hidden embedding space: steering LLMs to improve mental health assessment
- Sequences of Logits Reveal the Low Rank Structure of Language Models
- Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thought
- Do Stop Me Now: Detecting Boilerplate Responses with a Single Iteration
- Mapping Faithful Reasoning in Language Models
- Modeling Hierarchical Thinking in Large Reasoning Models
- Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
- Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training
- Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
- Agentic Reinforcement Learning for Search Misaligns Instruction-Tuning
- SARSteer: Safeguarding Large Audio-Language Models via Safe-Ablated Refusal Steering
- Who's Asking? Simulating Role-Based Questions for Conversational AI Evaluation
- To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models
- In-Distribution Steering: Balancing Control and Coherence in Language Model Generation
- Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment
- The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs
- RAID: Refusal-Aware and Integrated Decoding for Jailbreaking LLMs
- Keep Calm and Avoid Harmful Content: Concept Alignment and Latent Manipulation Towards Safer Answers
- Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
- Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research
- The Unintended Trade-off of AI Alignment:Balancing Hallucination Mitigation and Safety in LLMs
- PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector Algebra
- How Language Models Conflate Logical Validity with Plausibility: A Representational Analysis of Content Effects
- LatentBreak: Jailbreaking Large Language Models through Latent Space Feedback
- Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?
- Prototype-Based Dynamic Steering for Large Language Models
- Read the Scene, Not the Script: Outcome-Aware Safety for LLMs
- Activation Steering with a Feedback Controller
- MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
- A Granular Study of Safety Pretraining under Model Abliteration
- Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
- AbsTopK: Rethinking Sparse Autoencoders For Bidirectional Features
- ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
- Defeating Cerberus: Concept-Guided Privacy-Leakage Mitigation in Multimodal Language Models
- EasySteer: A Unified Framework for High-Performance and Extensible LLM Steering
- DiffuGuard: How Intrinsic Safety is Lost and Found in Diffusion Large Language Models
- Toward Preference-aligned Large Language Models via Residual-based Model Steering
- Beyond the Exploration-Exploitation Trade-off: A Hidden State Approach for LLM Reasoning in RLVR
- Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
- Enhancing LLM Steering through Sparse Autoencoder-Based Vector Refinement
- Train Once, Answer All: Many Pretraining Experiments for the Cost of One
- Knowledge distillation through geometry-aware representational alignment
- IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
- Detecting (Un)answerability in Large Language Models with Linear Directions
- The Rogue Scalpel: Activation Steering Compromises LLM Safety
- Active Attacks: Red-teaming LLMs via Adaptive Environments
- Hallucination reduction with CASAL: Contrastive Activation Steering For Amortized Learning
- Silent Tokens, Loud Effects: Padding in LLMs
- Compliance2LoRA: Personalizable On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters
- Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game
- Aligning Recommendations with User Popularity Preferences
- Steering When Necessary: Flexible Steering Large Language Models with Backtracking
- DISCO: Disentangled Communication Steering for Large Language Models
- SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection
- Representation-based Broad Hallucination Detectors Fail to Generalize Out of Distribution
- Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction
- LLM Jailbreak Detection for (Almost) Free!
- AdaptiveK Sparse Autoencoders: Dynamic Sparsity Allocation for Interpretable LLM Representations
- Programmable Cognitive Bias in Social Agents
- SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
- RepIt: Steering Language Models with Concept-Specific Refusal Vectors
- LLM Assertiveness can be Mechanistically Decomposed into Emotional and Logical Components
- Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks
- Unveiling the Latent Directions of Reflection in Large Language Models
- So let's replace this phrase with insult... Lessons learned from generation of toxic texts with LLMs
- Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint
- Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift
- Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
- MEUV: Achieving Fine-Grained Capability Activation in Large Language Models via Mutually Exclusive Unlock Vectors
- Unlearning That Lasts: Utility-Preserving, Robust, and Almost Irreversible Forgetting in LLMs
- From Indirect Object Identification to Syllogisms: Exploring Binary Mechanisms in Transformer Circuits
- ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
- Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
- Genre Controlled Music Generation via Activation Steering
- Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
- Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
- Safety Alignment Should Be Made More Than Just A Few Attention Heads
- Sycophancy as compositions of Atomic Psychometric Traits
- Learning to Steer: Input-dependent Steering for Multimodal LLMs
- Consiglieres in the Shadow: Understanding the Use of Uncensored Large Language Models in Cybercrimes
- Mitigating Jailbreaks with Intent-Aware LLMs
- Enhancing Supervised Composed Image Retrieval via Reasoning-Augmented Representation Engineering
- SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering
- Human-Alignment and Calibration of Inference-Time Uncertainty in Large Language Models
- VISOR: Visual Input-based Steering for Output Redirection in Vision-Language Models
- Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
- LoRA is All You Need for Safety Alignment of Reasoning LLMs
- Semantic Structure in Large Language Model Embeddings
- Don't Overthink It: A Survey of Efficient R1-style Large Reasoning Models
- Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders
- Multilingual Political Views of Large Language Models: Identification and Steering
- The Blessing and Curse of Dimensionality in Safety Alignment
- Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment
- SASFT: Sparse Autoencoder-guided Supervised Finetuning to Mitigate Unexpected Code-Switching in LLMs
- Security-by-Design for LLM-Based Code Generation: Leveraging Internal Representations for Concept-Driven Steering Mechanisms
- Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?
- When Models Examine Themselves: Vocabulary-Activation Correspondence in Self-Referential Processing
- GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt
- The Geometry of Harmfulness in LLMs through Subconcept Probing
- Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models
- Enhancing Cross-task Transfer of Large Language Models via Activation Steering
- Can We Predict Alignment Before Models Finish Thinking? Towards Monitoring Misaligned Reasoning Models
- LLMs Encode Harmfulness and Refusal Separately
- Reasoning-Finetuning Repurposes Latent Representations in Base Models
- AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety Basin
- BlueGlass: A Framework for Composite AI Safety
- Large Language Models Encode Semantics and Alignment in Linearly Separable Representations
- Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers
- Anthropomimetic Uncertainty: What Verbalized Uncertainty in Language Models is Missing
- Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates
- The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models
- Mechanistic Indicators of Understanding in Large Language Models
- Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents
- Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
- Where Do Reasoning Models Refuse?
- When Safety Becomes a Vulnerability: Exploiting LLM Alignment Homogeneity for Transferable Blocking in RAG
- Probing and Steering Evaluation Awareness of Language Models
- On Reasoning Strength Planning in Large Reasoning Models
- Failure by Interference: Language Models Make Balanced Parentheses Errors When Faulty Mechanisms Overshadow Sound Ones
- Linearly Decoding Refused Knowledge in Aligned Language Models
- STACK: Adversarial Attacks on LLM Safeguard Pipelines
- Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness
- Video Unlearning via Low-Rank Refusal Vector
- From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
- Safety-Aligned Weights Are Not Enough: Refusal-Teacher-Guided Finetuning Enhances Safety and Downstream Performance under Harmful Finetuning Attacks
- Persona Features Control Emergent Misalignment
- Improving LLM Reasoning through Interpretable Role-Playing Steering
- No Training Wheels: Steering Vectors for Bias Correction at Inference Time
- TwinBreak: Jailbreaking LLM Security Alignments based on Twin Prompts
- Understanding Reasoning in Thinking Language Models via Steering Vectors
- From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
- Latent Concept Disentanglement in Transformer-based Language Models
- Probing the Robustness of Large Language Models Safety to Latent Perturbations
- Exploring and Exploiting the Inherent Efficiency within Large Reasoning Models for Self-Guided Efficiency Enhancement
- LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning
- InverseScope: Scalable Activation Inversion for Interpreting Large Language Models
- Sysformer: Safeguarding Frozen Large Language Models with Adaptive System Prompts
- RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?
- Transferring Linear Features Across Language Models With Model Stitching
- Jailbreak Transferability Emerges from Shared Representations
- ContextBench: Modifying Contexts for Targeted Latent Activation
- Universal Jailbreak Suffixes Are Strong Attention Hijackers
- AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
- Improving Large Language Model Safety with Contrastive Representation Learning
- From Emergence to Control: Probing and Modulating Self-Reflection in Language Models
- Model Organisms for Emergent Misalignment
- Convergent Linear Representations of Emergent Misalignment
- Robustly Improving LLM Fairness in Realistic Settings via Interpretability
- Preserving Task-Relevant Information Under Linear Concept Removal
- SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation
- Evaluating Prompt-Driven Chinese Large Language Models: The Influence of Persona Assignment on Stereotypes and Safeguards
- Unveiling Language-Specific Features in Large Language Models via Sparse Autoencoders
- Line of Sight: On Linear Representations in VLLMs
- Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
- Steerable Chatbots: Exploring Personalization Control Interfaces via LLM Activation Steering
- From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
- IF-GUIDE: Influence Function-Guided Detoxification of LLMs
- SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs
- Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models
- Beyond Multiple Choice: Evaluating Steering Vectors for Summarization
- Circuit Stability Characterizes Language Model Generalization
- Learning Safety Constraints for Large Language Models
- Quiet Feature Learning in Algorithmic Tasks
- COSMIC: Generalized Refusal Direction Identification in LLM Activations
- Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
- Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
- SafeScientist: Toward Risk-Aware Scientific Discoveries by LLM Agents
- Does Machine Unlearning Truly Remove Knowledge?
- SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models
- MCP Safety Training: Learning to Refuse Falsely Benign MCP Exploits using Improved Preference Alignment
- Understanding Refusal in Language Models with Sparse Autoencoders
- Derailing Non-Answers via Logit Suppression at Output Subspace Boundaries in RLHF-Aligned Language Models
- Mitigating Overthinking in Large Reasoning Models via Manifold Steering
- Precise In-Parameter Concept Erasure in Large Language Models
- Understanding (Un)Reliability of Steering Vectors in Language Models
- Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing
- From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs
- Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders
- Paths Not Taken: Understanding and Mending the Multilingual Factual Recall Pipeline
- ExPLAIND: Unifying Model, Data, and Training Attribution to Study Model Behavior
- Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models
- An Embarrassingly Simple Defense Against LLM Abliteration Attacks
- From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test
- Do BERT-Like Bidirectional Models Still Perform Better on Text Classification in the Era of LLMs?
- Refusal Direction is Universal Across Safety-Aligned Languages
- Robustifying Vision-Language Models via Dynamic Token Reweighting
- SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models
- EnSToM: Enhancing Dialogue Systems with Entropy-Scaled Steering Vectors for Topic Maintenance
- Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models
- When Do LLMs Admit Their Mistakes? Understanding The Role Of Model Belief In Retraction
- The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness
- Safety Degradation in AI Agents
- SoK: Intent-Oriented Systematization of Multi-Turn LLM Jailbreaks
- Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case Study
- Language Models Are Capable of Metacognitive Monitoring and Control of Their Internal Activations
- Improving Multilingual Language Models by Aligning Representations through Steering
- Self-Destructive Language Model
- Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
- Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
- A False Average: Chain-of-Thought Monitors Collapse Where They Are the Only Defense
- Noise Injection Systemically Degrades Large Language Model Safety Guardrails
- Layered Unlearning for Adversarial Relearning
- Can LLMs Reliably Self-Report Adversarial Prefills, and How?
- FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
- Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models
- LangFIR: Discovering Sparse Language-Specific Features from Monolingual Data for Language Steering
- On the Limitations of Steering in Language Model Alignment
- TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
- Measuring and Eliminating Refusals in Military Large Language Models
- Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks
- Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure
- Fast Multi-dimensional Refusal Subspaces via RFM-AGOP
- When Behavioral Safety Evaluation Fails: A Representation-Level Perspective
- Valence-Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control
- The Value Axis: Language Models Encode Whether They're on the Right Track
- Beyond Refusal: A Same-Lineage Study of Aligned and Abliterated LLMs for Vulnerability Analysis
- Positive Alignment: Artificial Intelligence for Human Flourishing
- SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
- LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations
- Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry
- Targeted Neuron Modulation via Contrastive Pair Search
- Fair outputs, Biased Internals: Causal Potency and Asymmetry of Latent Bias in LLMs for High-Stakes Decisions
- Estimating Tail Risks in Language Model Output Distributions
- Semantic Structure of Feature Space in Large Language Models
- What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations?
- Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection
- HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models
- Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning
- Subliminal Learning Is Steering Vector Distillation
- Steered LLM Activations are Non-Surjective
- Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types
- The Information Geometry of Softmax: Probing and Steering
- There Is More to Refusal in Large Language Models than a Single Direction
- Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration
- Adaptive Helpfulness-Harmlessness Alignment with Preference Vectors
- Reasoning Models Know What's Important, and Encode It in Their Activations
- Surgical Repair of Insecure Code Generation in LLMs
- What's the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering
- When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models
- A Theory of Conditional Collapse under Low-Rank Weight-Space Ablations: I. The Single-Block Theory and Synthetic Validation
- DataRx: Missingness-Aware Sampling for Safer Large Language Model Task-Specific Fine-Tuning
- Test, then Route: How Language Models Execute In-Context Conditional Rules Across Models and Languages
- Language Models Are Implicitly Continuous
- Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
- Activation Patching for Interpretable Steering in Music Generation
- Identifying and Manipulating Personality Traits in LLMs Through Activation Engineering
- FairSteer: Inference Time Debiasing for LLMs with Dynamic Activation Steering
- The Geometry of Self-Verification in a Task-Specific Reasoning Model
- Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
- CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits
- Detecting Safety Training Modification in Language Models via Activation Analysis
- MMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration
- A Mechanistic Analysis of Gender Sensitivity in Dense Retrieval Models
- Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment
- Safe Evolution with Circuit Anchors
- RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
- AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
- The Structural Safety Generalization Problem
- Deceptive Automated Interpretability: Language Models Coordinating to Fool Oversight Systems
- On the Effectiveness and Generalization of Race Representations for Debiasing High-Stakes Decisions
Discussions
- Refusal in language models is mediated by a single direction [hn, 209 points, 44 comments]
- Refusal in Language Models Is Mediated by a Single Direction [hn, 118 points, 45 comments]
- For example, Arditi et al. (arxiv.org/abs/2406.11717) argued that refusal is mediated by a single "direction," or linear subspace, in many language models. But when we applied their method on a RWKV m [bsky, 4 points, 1 comments]
- My take: models are capable of learning some patterns, and not others. But, without a causal experiment, we don't know a priori know what these patterns are. This is my favorite recent example that e [bsky, 2 points, 0 comments]
- Training LLMs includes teaching them to sometimes respond "I'm sorry, but I can't answer that". AI research calls this "refusal" and it is one of many separable proto-concepts in these systems. This A [bsky, 2 points, 1 comments]
- Yes, the technique they used is the same technique that heretic uses! (From https://arxiv.org/abs/2406.11717) [bsky, 1 points, 0 comments]
- Refusal in language models is mediated by a single direction (arxiv.org) Main Link | Discussion [bsky, 1 points, 0 comments]
- Refusal in Language Models Is Mediated by a Single Direction https://arxiv.org/abs/2406.11717 [comments] [70 points] [bsky, 0 points, 0 comments]
- I *suspect* you run into an issue similar to the one described in part 5 of "Refusal in LMs is mediated by a single direction" arxiv.org/abs/2406.11717 where "refuse injected commands" becomes part of [bsky, 0 points, 1 comments]
- Re: Jailbreaks I suspect mechanistic analysis of various jailbreaks will give us a better idea of how to prevent them. I feel like there ultimately isn't a fix for something like this prefix attack th [bsky, 0 points, 1 comments]
- Models refusing requests through one linear direction in activation space? That's surprisingly simple – and probably why jailbreaks work so well. https://arxiv.org/abs/2406.11717 [bsky, 0 points, 1 comments]
- https://arxiv.org/abs/2406.11717 大規模言語モデルの拒否挙動は単一方向で制御可能であることを発見しました。 この方向を削除すると拒否がなくなり、追加すると無害な指示も拒否します。 安全対策の脆弱性を示唆し、モデル内部理解の重要性を強調しています。 [bsky, 0 points, 0 comments]
- Refusal in Language Models Is Mediated by a Single Direction https:// arxiv.org/abs/2406.11717 # arxiv [mastodon, 0 points, 0 comments]
- Refusal in Language Models Is Mediated by a Single Direction https://arxiv.org/abs/2406.11717 [bsky, 0 points, 0 comments]
- Refusal in Language Models Is Mediated by a Single Direction https://arxiv.org/abs/2406.11717 (https://news.ycombinator.com/item?id=47986136) [bsky, 0 points, 0 comments]
- Refusal in Language Models Is Mediated by a Single Direction https://arxiv.org/abs/2406.11717 (https://news.ycombinator.com/item?id=47986136) [bsky, 0 points, 0 comments]
- Refusal in Language Models Is Mediated by a Single Direction [bsky, 0 points, 0 comments]
- Refusal in Language Models Is Mediated by a Single Direction #HackerNews https://arxiv.org/abs/2406.11717 [bsky, 0 points, 0 comments]
- Refusal in Language Models Is Mediated by a Single Direction https://arxiv.org/abs/2406.11717 https://news.ycombinator.com/item?id=47986136 [bsky, 0 points, 0 comments]
- Refusal in Language Models Is Mediated by a Single Direction arxiv.org/pdf/2406.11717 [bsky, 0 points, 0 comments]
Related