vix.ing · top · new · best · stats · spec

Dongkyu Derek Cho

  1. Breaking the Safety-Capability Tradeoff: Reinforcement Learning with Verifiable Rewards Maintains Safety Guardrails in LLMs
    2025/11/26 by Dongkyu Derek Cho, Huan Song, Cho, Dongkyu Derek +15 · 1 citation
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI)