vix.ing · top · new · best · stats · spec

Shi, Dongsheng

  1. Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks
    2025/01/18 by Yi Xin, Yi, Xin, Ting Li +8 · 2 citations
    Computer Science · #Adversarial Robustness in Machine Learning
  2. Unified Defense for Large Language Models against Jailbreak and Fine-Tuning Attacks in Education
    2025/11/18 by Yi, Xin, Li, Yue, Shi, Dongsheng +3 · 1 citation
    Computer Science · #Adversarial Robustness in Machine Learning #Topic Modeling #Explainable Artificial Intelligence (XAI)