vix.ing · top · new · best · stats · spec

Qi, Xiangyu

  1. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
    2023/10/05 by Xiangyu Qi, Yi Zeng, Qi, Xiangyu +11 · 164 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  2. Visual Adversarial Examples Jailbreak Aligned Large Language Models
    2023/06/22 by Xiangyu Qi, Kaixuan Huang, Qi, Xiangyu +8 · 69 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  3. Safety Alignment Should Be Made More Than Just a Few Tokens Deep
    2024/06/10 by Qi, Xiangyu, Panda, Ashwinee, Lyu, Kaifeng +5 · 93 citations
    #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
  4. SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
    2024/06/20 by Tinghao Xie, Xiangyu Qi, Xie, Tinghao +29 · 39 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Natural Language Processing Techniques #Software Reliability and Analysis Research #Topic Modeling
  5. Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
    2024/02/07 by Boyi Wei, Kaixuan Huang, Wei, Boyi +15 · 33 citations
    Decision Sciences · Engineering · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Fatigue and fracture mechanics #Machine Learning (cs.LG) #Risk and Safety Analysis
  6. Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
    2024/05/30 by Xiong Chen, Xiong, Chen, Xiangyu Qi +5 · 11 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Cryptography and Data Security #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
  7. Lottery Ticket Adaptation: Mitigating Destructive Interference in LLMs
    2024/06/24 by Ashwinee Panda, Berivan Isik, Panda, Ashwinee +9 · 8 citations
    Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Auction Theory and Applications #Computation and Language (cs.CL) #Digital Rights Management and Security #FOS: Computer and information sciences
  8. On Evaluating the Durability of Safeguards for Open-Weight LLMs
    2024/12/10 by Xiangyu Qi, Boyi Wei, Qi, Xiangyu +17 · 2 voices · 7 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #cs.AI #cs.CR
  9. Knowledge Enhanced Machine Learning Pipeline against Diverse Adversarial Attacks
    2021/06/11 by Gürel, Nezihe Merve, Qi, Xiangyu, Rimanic, Luka +2 · 3 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG)
  10. Towards Practical Deployment-Stage Backdoor Attack on Deep Neural Networks
    2021/11/25 by Qi, Xiangyu, Xie, Tinghao, Pan, Ruizhe +3 · 3 citations
    #Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
  11. Towards A Proactive ML Approach for Detecting Backdoor Poison Samples
    2022/05/26 by Qi, Xiangyu, Xie, Tinghao, Wang, Jiachen T. +3 · 3 citations
    #Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  12. Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
    2024/02/22 by Wang, Jiongxiao, Li, Jiazhao, Li, Yiquan +7 · 4 citations
    #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
  13. Subnet Replacement: Deployment-stage backdoor attack against deep neural networks in gray-box setting
    2021/07/15 by Xiangyu Qi, Qi, Xiangyu, Zhu Jifeng +5 · 1 citation
    Computer Science · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  14. Uncovering Adversarial Risks of Test-Time Adaptation
    2023/01/29 by Tong Wu, Wu, Tong, Feiran Jia +11 · 1 citation
    Computer Science · Biochemistry, Genetics and Molecular Biology · Medicine · #Anomaly Detection Techniques and Applications #Metabolomics and Mass Spectrometry Studies #Data-Driven Disease Surveillance
  15. BaDExpert: Extracting Backdoor Functionality for Accurate Backdoor Input Detection
    2023/08/23 by Tinghao Xie, Xiangyu Qi, Xie, Tinghao +9 · 1 citation
    Computer Science · #Adversarial Robustness in Machine Learning
  16. AI Risk Management Should Incorporate Both Safety and Security
    2024/05/29 by Xiangyu Qi, Yangsibo Huang, Qi, Xiangyu +47 · 1 citation
    Medicine · Social Sciences · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Healthcare and Education #Cryptography and Security (cs.CR) #Ethics and Social Impacts of AI #FOS: Computer and information sciences