Sicheng Zhu
- AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models
2023/10/23 by Sicheng Zhu, Ruiyi Zhang, Zhu, Sicheng +15 · 23 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- On the Possibilities of AI-Generated Text Detection
2023/04/10 by Souradip Chakraborty, Amrit Singh Bedi, Chakraborty, Souradip +9 · 11 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques
- Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
2024/09/01 by Bang An, Sicheng Zhu, An, Bang +9 · 1 voice · 6 citations
Computer Science · #Computation and Language (cs.CL) #Computers and Society (cs.CY) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.CR #cs.CY #cs.LG
- Can Watermarking Large Language Models Prevent Copyrighted Text Generation and Hide Training Data?
2024/07/24 by Michael-Andrei Panaitescu-Liess, Zora Che, Panaitescu-Liess, Michael-Andrei +15 · 2 voices · 4 citations
#cs.LG
- AdvPrefix: An Objective for Nuanced LLM Jailbreaks
2024/12/13 by Sicheng Zhu, Brandon Amos, Zhu, Sicheng +7 · 7 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #Digital and Cyber Forensics #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Sex‐Specific Classification of Drug‐Induced Torsade de Pointes Susceptibility Using Cardiac Simulations and Machine Learning
2021/03/26 by Alex Fogli Iseppe, Haibo Ni, Sicheng Zhu +9 · 26 citations
Medicine · Biochemistry, Genetics and Molecular Biology · #Cardiac electrophysiology and arrhythmias #Ion channel regulation and function #Receptor Mechanisms and Signaling
- Learning Adversarially Robust Representations via Worst-Case Mutual Information Maximization
2020/02/26 by Sicheng Zhu, Zhu, Sicheng, Xiao Zhang +3 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Cryptography and Security (cs.CR) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Information Theory (cs.IT) #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- GPT-Red: Automated Red Teaming via Self-Play at Scale
2026/07/28 by Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal +15
Computer Science · #cs.AI #cs.CL #cs.CR #cs.LG