Dongkyu Derek Cho
- Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs
2025/11/26 by Dongkyu Derek Cho, Huan Song, Cho, Dongkyu Derek +15 · 1 citation
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI)