Javier Rando
- Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI
2024/06/17 by Robert Hönig, Javier Rando, Hönig, Robert +5 · 38 voices · 7 citations
Computer Science · #Adversarial Robustness in Machine Learning #cs.CR
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
2025/10/08 by Alexandra Souly, Souly, Alexandra, Javier Rando +25 · 34 voices · 18 citations
Medicine · #Medical Imaging and Pathology Studies
- Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
2023/07/27 by Stephen Casper, Casper, Stephen, Xander Davies +65 · 3 voices · 86 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Reliability and Analysis Research #cs.AI #cs.CL #cs.LG
- Persistent Pre-Training Poisoning of LLMs
2024/10/17 by Yiming Zhang, Zhang, Yiming, Javier Rando +13 · 3 voices · 20 citations
#cs.CR #cs.AI
- Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation
2023/11/06 by Rusheb Shah, Quentin Feuillade--Montixi, Shah, Rusheb +9 · 1 voice · 30 citations
Computer Science · Social Sciences · #Ethics and Social Impacts of AI #Persona Design and Applications #cs.AI #cs.CL #cs.LG
- Red-Teaming the Stable Diffusion Safety Filter
2022/10/03 by Javier Rando, Daniel Paleka, Rando, Javier +6 · 26 citations
Computer Science · #Generative Adversarial Networks and Image Synthesis #Digital Media Forensic Detection #Adversarial Robustness in Machine Learning
- Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations
2024/11/15 by Jianfeng Chi, Ujjwal Karn, Chi, Jianfeng +17 · 37 citations
Earth and Planetary Sciences · #3D Surveying and Cultural Heritage #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- An Adversarial Perspective on Machine Unlearning for AI Safety
2024/09/26 by Jakub Łucki, Łucki, Jakub, Boyi Wei +9 · 3 voices · 19 citations
Computer Science · #Adversarial Robustness in Machine Learning #cs.AI #cs.CL #cs.CR #cs.LG
- Universal Jailbreak Backdoors from Poisoned Human Feedback
2023/11/24 by Javier Rando, Rando, Javier, Florian Tramèr +1 · 16 citations
Computer Science · #Adversarial Robustness in Machine Learning #Hate Speech and Cyberbullying Detection
- Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
2024/04/22 by Javier Rando, Rando, Javier, Francesco Croce +11 · 4 citations
Computer Science · Economics, Econometrics and Finance · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Law, AI, and Intellectual Property #Law, Economics, and Judicial Systems #Machine Learning (cs.LG)
- Measuring Non-Adversarial Reproduction of Training Data in Large Language Models
2024/11/15 by Michael Aerni, Javier Rando, Aerni, Michael +9 · 2 voices · 2 citations
Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.LG
- Gradient-based Jailbreak Images for Multimodal Fusion Models
2024/10/04 by Javier Rando, Rando, Javier, Hannah Korevaar +7 · 2 citations
Engineering · Medicine · #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Medical Imaging Techniques and Applications #Medical Imaging and Analysis
- Exploring Adversarial Attacks and Defenses in Vision Transformers trained with DINO
2022/06/14 by Javier Rando, Rando, Javier, Nasib Naimi +5 · 1 citation
Computer Science · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
2025/03/03 by Nicholas Carlini, Carlini, Nicholas, Javier Rando +7 · 3 citations
Computer Science · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Security and Verification in Computing