RAJ-PGA: Reasoning-Activated Jailbreak and Principle-Guided Alignment Framework for Large Reasoning Models
2025/08/18 by Jianhao Chen, Chen, Jianhao, M. Xu +11
Computer Science · Engineering · #Fuzzy Logic and Control Systems #Software Reliability and Analysis Research #Infrastructure Maintenance and Monitoring
paper · pdf · doi:10.48550/arxiv.2508.12897
Abstract
Large Reasoning Models (LRMs) face a distinct safety vulnerability: their internal reasoning chains may generate harmful content even when the final output appears benign. To address this overlooked risk, we first propose a novel attack paradigm, Reasoning-Activated Jailbreak (RAJ) via Concretization, which demonstrates that refining malicious prompts to be more specific can trigger step-by-step logical reasoning that overrides the model's safety protocols. To systematically mitigate this vulnerability, we further develop a scalable framework for constructing high-quality safety alignment datasets. This framework first leverages the RAJ attack to elicit challenging harmful reasoning chains from LRMs, then transforms these high-risk traces into safe, constructive, and educational responses through a tailored Principle-Guided Alignment (PGA) mechanism. Then, we introduce the PGA dataset, a verified alignment dataset containing 3,989 samples using our proposed method. Extensive experiments show that fine-tuning LRMs with PGA dataset significantly enhances model safety, achieving up to a 29.5% improvement in defense success rates across multiple jailbreak benchmarks. Critically, our approach not only defends against sophisticated reasoning-based attacks but also preserves, even enhances, the model's general reasoning capabilities. This work provides a scalable and effective pathway for safety alignment in reasoning-intensive AI systems, addressing the core trade-off between safety and functional performance.
Citations
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Qwen3Guard Technical Report
- UnsafeChain: Enhancing Reasoning Model Safety via Hard Cases
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Three Minds, One Legend: Jailbreak Large Reasoning Model with Adaptive Stacked Ciphers
- Qwen3 Technical Report
- Practical Reasoning Interruption Attacks on Reasoning Large Language Models
- 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
- STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
- A Survey of Efficient Reasoning for Large Reasoning Models: Language, Multimodality, and Beyond
- Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
- From System 1 to System 2: A Survey of Reasoning Large Language Models
- The Hidden Risks of Large Reasoning Models: A Safety Assessment of R1
- SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
- STAIR: Improving Safety Alignment with Introspective Reasoning
- o3-mini vs DeepSeek-R1: Which One is Safer?
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- A Taxonomy of Systemic Risks from General-Purpose AI
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
- PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
- WildChat: 1M ChatGPT Interaction Logs in the Wild
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
- A StrongREJECT for Empty Jailbreaks
- Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
- Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Jailbroken: How Does LLM Safety Training Fail?
- Let's Verify Step by Step
- Affective Coherence Monitoring for Transformer-Based Language Models
- Why Should Adversarial Perturbations be Imperceptible? Rethink the Research Paradigm in Adversarial NLP
- Training language models to follow instructions with human feedback
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- Evaluating Large Language Models Trained on Code
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- OpenAI o1 System Card
Related