vix.ing · top · new · best · stats · spec

Zeming Wei

  1. Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
    2023/10/10 by Zeming Wei, Wei, Zeming, Yifei Wang +4 · 82 citations
    Computer Science · #Topic Modeling #Adversarial Robustness in Machine Learning #Hate Speech and Cyberbullying Detection
  2. Jatmo: Prompt Injection Defense by Task-Specific Finetuning
    2023/12/29 by Julien Piet, Maha Al-Rashed, Piet, Julien +13 · 27 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Topic Modeling #Natural Language Processing Techniques
  3. A Theoretical Understanding of Self-Correction through In-context Alignment
    2024/05/28 by Yifei Wang, Wang, Yifei, Yuyang Wu +7 · 12 citations
    Social Sciences · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Multimedia Communication and Technology
  4. Boosting Jailbreak Attack with Momentum
    2024/05/02 by Yihao Zhang, Zhang, Yihao, Zeming Wei +1 · 10 citations
    Computer Science · #Advanced Malware Detection Techniques #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Network Security and Intrusion Detection #Optimization and Control (math.OC)
  5. CFA: Class-wise Calibrated Fair Adversarial Training
    2023/03/25 by Zeming Wei, Yifei Wang, Wei, Zeming +5 · 7 citations
    Computer Science · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  6. Sharpness-Aware Minimization Alone can Improve Adversarial Robustness
    2023/05/09 by Zeming Wei, Wei, Zeming, Jingyu Zhu +3 · 5 citations
    Computer Science · Engineering · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Ferroelectric and Negative Capacitance Devices #Machine Learning (cs.LG)
  7. Fight Back Against Jailbreaking via Prompt Adversarial Tuning
    2024/02/09 by Yichuan Mo, Mo, Yichuan, Yuji Wang +5 · 6 citations
    Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #Digital and Cyber Forensics #FOS: Computer and information sciences #Law in Society and Culture #Machine Learning (cs.LG)
  8. Exploring the Robustness of In-Context Learning with Noisy Labels
    2024/04/28 by Cheng Chen, Xinzhi Yu, Cheng, Chen +10 · 2 citations
    Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning and Data Classification #Optimization and Control (math.OC) #Water Systems and Optimization #Wireless Sensor Networks and IoT
  9. Weighted Automata Extraction and Explanation of Recurrent Neural Networks for Natural Language Tasks
    2023/06/24 by Zeming Wei, Wei, Zeming, Xiyue Zhang +5 · 1 citation
    Computer Science · Engineering · Materials Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #FOS: Computer and information sciences #Ferroelectric and Negative Capacitance Devices #Machine Learning (cs.LG) #Machine Learning in Materials Science
  10. Architecture Matters: Uncovering Implicit Mechanisms in Graph Contrastive Learning
    2023/11/05 by Xiaojun Guo, Guo, Xiaojun, Yifei Wang +5 · 1 citation
    Computer Science · Health Professions · #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Health Literacy and Information Accessibility #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  11. 3DAffordSplat: Efficient Affordance Reasoning with 3D Gaussians
    2025/04/15 by Zeming Wei, Wei, Zeming, Junyi Lin +11 · 3 citations
    Engineering · #Robot Manipulation and Learning #3D Shape Modeling and Analysis #Robotics and Sensor-Based Localization
  12. Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
    2025/05/21 by Taiye Chen, Zeming Wei, Chen, Taiye +5 · 2 citations
    Computer Science · #Advanced Malware Detection Techniques #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #Digital and Cyber Forensics #FOS: Computer and information sciences #Machine Learning (cs.LG)
  13. Advancing LLM Safe Alignment with Safety Representation Ranking
    2025/05/21 by Tianqi Du, Du, Tianqi, Zeming Wei +7 · 3 citations
    Business, Management and Accounting · Decision Sciences · Health Professions · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Occupational Health and Safety Research #Quality and Management Systems #Risk and Safety Analysis
  14. Towards the Worst-case Robustness of Large Language Models
    2025/01/31 by Huanran Chen, Chen, Huanran, Yinpeng Dong +7 · 1 citation
    Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
  15. Hyperball May Not Be a Free Lunch
    2026/07/24 by Yihao Xiao, Jialong Sun, Zitian Gao +5
    #cs.LG #cs.AI