Alexander Robey
- SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
2023/10/05 by Alexander Robey, Eric Wong, Robey, Alexander +5 · 2 voices · 80 citations
Computer Science · #Topic Modeling #Adversarial Robustness in Machine Learning #Natural Language Processing Techniques
- Jailbreaking Black Box Large Language Models in Twenty Queries
2023/10/12 by Patrick Chao, Chao, Patrick, Alexander Robey +9 · 217 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
2024/03/28 by Patrick Chao, Chao, Patrick, Edoardo Debenedetti +21 · 90 citations
Computer Science · #Authorship Attribution and Profiling #Cryptography and Security (cs.CR) #Digital and Cyber Forensics #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques
- Efficient and Accurate Estimation of Lipschitz Constants for Deep Neural Networks
2019/06/11 by Mahyar Fazlyab, Alexander Robey, Fazlyab, Mahyar +7 · 17 citations
Computer Science · #Adversarial Robustness in Machine Learning #Machine Learning and Algorithms #Domain Adaptation and Few-Shot Learning
- A Safe Harbor for AI Evaluation and Red Teaming
2024/03/07 by Shayne Longpre, Sayash Kapoor, Longpre, Shayne +43 · 1 voice · 15 citations
Computer Science · #Explainable Artificial Intelligence (XAI)
- Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing
2024/02/25 by Jiabao Ji, Ji, Jiabao, Bairu Hou +12 · 10 citations
Computer Science · #Adversarial Robustness in Machine Learning #Hate Speech and Cyberbullying Detection #Digital and Cyber Forensics
- Model-Based Domain Generalization
2021/02/23 by Alexander Robey, George J. Pappas, Robey, Alexander +3 · 4 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Multimodal Machine Learning Applications #Machine Learning and Algorithms
- Learning Hybrid Control Barrier Functions from Data
2020/11/08 by Lars Lindemann, Haimin Hu, Lindemann, Lars +11 · 4 citations
Engineering · #Advanced Control Systems Optimization #Fault Detection and Control Systems #Smart Grid Security and Resilience
- Safety Pretraining: Toward the Next Generation of Safe AI
2025/04/23 by Pratyush Maini, Sachin Goyal, Maini, Pratyush +18 · 3 voices · 8 citations
Health Professions · #cs.LG
- Jailbreaking LLM-Controlled Robots
2024/10/17 by Alexander Robey, Zachary Ravichandran, Robey, Alexander +7 · 6 voices
Computer Science · Engineering · #Advanced Malware Detection Techniques #Modular Robots and Swarm Intelligence #Digital and Cyber Forensics
- Toward Certified Robustness Against Real-World Distribution Shifts
2022/06/08 by Haoze Wu, Teruhiro Tagomori, Wu, Haoze +15 · 3 citations
Computer Science · Physics and Astronomy · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Logic in Computer Science (cs.LO) #Machine Learning (cs.LG) #Model Reduction and Neural Networks
- Safety Guardrails for LLM-Enabled Robots
2025/03/10 by Zachary Ravichandran, Ravichandran, Zachary, Alexander Robey +7 · 7 citations
Engineering · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Real-time simulation and control systems #Robotics (cs.RO) #Transportation Safety and Impact Analysis
- Adversarial Robustness with Semi-Infinite Constrained Learning
2021/10/29 by Alexander Robey, Luiz F. O. Chamon, Robey, Alexander +7 · 2 citations
Computer Science · Engineering · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #FOS: Computer and information sciences #Fault Detection and Control Systems #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Chordal Sparsity for Lipschitz Constant Estimation of Deep Neural Networks
2022/04/02 by Anton Xue, Lars Lindemann, Xue, Anton +9 · 2 citations
Engineering · Computer Science · #Sparse and Compressive Sensing Techniques #Adversarial Robustness in Machine Learning #Machine Learning and Algorithms
- Optimal Algorithms for Submodular Maximization with Distributed Constraints
2019/09/30 by Alexander Robey, Arman Adibi, Robey, Alexander +7 · 1 citation
Computer Science · #Complexity and Algorithms in Graphs #Cryptography and Data Security #Data Structures and Algorithms (cs.DS) #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Optimization and Control (math.OC) #Optimization and Search Problems
- Provable tradeoffs in adversarially robust classification
2020/06/09 by Edgar Dobriban, Hamed Hassani, Dobriban, Edgar +5 · 1 citation
Computer Science · Decision Sciences · Mathematics · #Advanced Statistical Methods and Models #Advanced Statistical Process Monitoring #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- On the Sample Complexity of Stability Constrained Imitation Learning
2021/02/18 by Stephen Tu, Tu, Stephen, Alexander Robey +5 · 2 citations
Computer Science · Engineering · Physics and Astronomy · #Reinforcement Learning in Robotics #Robot Manipulation and Learning #Model Reduction and Neural Networks
- Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks
2025/02/28 by Hanjiang Hu, Alexander Robey, Hu, Hanjiang +3 · 4 citations
Psychology · #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Social Robot Interaction and HRI
- Existing Large Language Model Unlearning Evaluations Are Inconclusive
2025/05/31 by Zhili Feng, Feng, Zhili, Ye Xu +13 · 4 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling
- Adversarial Training Should Be Cast as a Non-Zero-Sum Game
2023/06/19 by Alexander Robey, Robey, Alexander, Fabian Latorre +7 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Optimization and Control (math.OC)
- Transferable Adversarial Attacks on Black-Box Vision-Language Models
2025/05/02 by Kai Hu, Weichen Yu, Hu, Kai +13 · 3 citations
Computer Science · #Adversarial Robustness in Machine Learning
- Embodied AI: Emerging Risks and Opportunities for Policy Action
2025/08/28 by Jared Perlo, Alexander Robey, Perlo, Jared +7 · 1 voice · 2 citations
Computer Science · Psychology · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Human-Automation Interaction and Safety #cs.AI #cs.CY #cs.RO