Qi, Xiangyu
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
2023/10/05 by Xiangyu Qi, Yi Zeng, Qi, Xiangyu +11 · 164 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Visual Adversarial Examples Jailbreak Aligned Large Language Models
2023/06/22 by Xiangyu Qi, Kaixuan Huang, Qi, Xiangyu +8 · 69 citations
Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Safety Alignment Should Be Made More Than Just a Few Tokens Deep
2024/06/10 by Qi, Xiangyu, Panda, Ashwinee, Lyu, Kaifeng +5 · 93 citations
#Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
- SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
2024/06/20 by Tinghao Xie, Xiangyu Qi, Xie, Tinghao +29 · 39 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Natural Language Processing Techniques #Software Reliability and Analysis Research #Topic Modeling
- Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
2024/02/07 by Boyi Wei, Kaixuan Huang, Wei, Boyi +15 · 33 citations
Decision Sciences · Engineering · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Fatigue and fracture mechanics #Machine Learning (cs.LG) #Risk and Safety Analysis
- Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
2024/05/30 by Xiong Chen, Xiong, Chen, Xiangyu Qi +5 · 11 citations
Computer Science · #Adversarial Robustness in Machine Learning #Cryptography and Data Security #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
- Lottery Ticket Adaptation: Mitigating Destructive Interference in LLMs
2024/06/24 by Ashwinee Panda, Berivan Isik, Panda, Ashwinee +9 · 8 citations
Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Auction Theory and Applications #Computation and Language (cs.CL) #Digital Rights Management and Security #FOS: Computer and information sciences
- On Evaluating the Durability of Safeguards for Open-Weight LLMs
2024/12/10 by Xiangyu Qi, Boyi Wei, Qi, Xiangyu +17 · 2 voices · 7 citations
Computer Science · #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #cs.AI #cs.CR
- Knowledge Enhanced Machine Learning Pipeline against Diverse Adversarial Attacks
2021/06/11 by Gürel, Nezihe Merve, Qi, Xiangyu, Rimanic, Luka +2 · 3 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Towards Practical Deployment-Stage Backdoor Attack on Deep Neural Networks
2021/11/25 by Qi, Xiangyu, Xie, Tinghao, Pan, Ruizhe +3 · 3 citations
#Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
- Towards A Proactive ML Approach for Detecting Backdoor Poison Samples
2022/05/26 by Qi, Xiangyu, Xie, Tinghao, Wang, Jiachen T. +3 · 3 citations
#Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment
2024/02/22 by Wang, Jiongxiao, Li, Jiazhao, Li, Yiquan +7 · 4 citations
#Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
- Subnet Replacement: Deployment-stage backdoor attack against deep neural networks in gray-box setting
2021/07/15 by Xiangyu Qi, Qi, Xiangyu, Zhu Jifeng +5 · 1 citation
Computer Science · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Uncovering Adversarial Risks of Test-Time Adaptation
2023/01/29 by Tong Wu, Wu, Tong, Feiran Jia +11 · 1 citation
Computer Science · Biochemistry, Genetics and Molecular Biology · Medicine · #Anomaly Detection Techniques and Applications #Metabolomics and Mass Spectrometry Studies #Data-Driven Disease Surveillance
- BaDExpert: Extracting Backdoor Functionality for Accurate Backdoor Input Detection
2023/08/23 by Tinghao Xie, Xiangyu Qi, Xie, Tinghao +9 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning
- AI Risk Management Should Incorporate Both Safety and Security
2024/05/29 by Xiangyu Qi, Yangsibo Huang, Qi, Xiangyu +47 · 1 citation
Medicine · Social Sciences · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Healthcare and Education #Cryptography and Security (cs.CR) #Ethics and Social Impacts of AI #FOS: Computer and information sciences