vix.ing · top · new · best · stats · spec

Javier Rando

  1. Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI
    2024/06/17 by Robert Hönig, Javier Rando, Hönig, Robert +5 · 38 voices · 7 citations
    Computer Science · #Adversarial Robustness in Machine Learning #cs.CR
  2. Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
    2025/10/08 by Alexandra Souly, Souly, Alexandra, Javier Rando +25 · 34 voices · 18 citations
    Medicine · #Medical Imaging and Pathology Studies
  3. Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
    2023/07/27 by Stephen Casper, Casper, Stephen, Xander Davies +65 · 3 voices · 86 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Reliability and Analysis Research #cs.AI #cs.CL #cs.LG
  4. Persistent Pre-Training Poisoning of LLMs
    2024/10/17 by Yiming Zhang, Zhang, Yiming, Javier Rando +13 · 3 voices · 20 citations
    #cs.CR #cs.AI
  5. Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation
    2023/11/06 by Rusheb Shah, Quentin Feuillade--Montixi, Shah, Rusheb +9 · 1 voice · 30 citations
    Computer Science · Social Sciences · #Ethics and Social Impacts of AI #Persona Design and Applications #cs.AI #cs.CL #cs.LG
  6. Red-Teaming the Stable Diffusion Safety Filter
    2022/10/03 by Javier Rando, Daniel Paleka, Rando, Javier +6 · 26 citations
    Computer Science · #Generative Adversarial Networks and Image Synthesis #Digital Media Forensic Detection #Adversarial Robustness in Machine Learning
  7. Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations
    2024/11/15 by Jianfeng Chi, Ujjwal Karn, Chi, Jianfeng +17 · 37 citations
    Earth and Planetary Sciences · #3D Surveying and Cultural Heritage #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  8. An Adversarial Perspective on Machine Unlearning for AI Safety
    2024/09/26 by Jakub Łucki, Łucki, Jakub, Boyi Wei +9 · 3 voices · 19 citations
    Computer Science · #Adversarial Robustness in Machine Learning #cs.AI #cs.CL #cs.CR #cs.LG
  9. Universal Jailbreak Backdoors from Poisoned Human Feedback
    2023/11/24 by Javier Rando, Rando, Javier, Florian Tramèr +1 · 16 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Hate Speech and Cyberbullying Detection
  10. Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
    2024/04/22 by Javier Rando, Rando, Javier, Francesco Croce +11 · 4 citations
    Computer Science · Economics, Econometrics and Finance · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Law, AI, and Intellectual Property #Law, Economics, and Judicial Systems #Machine Learning (cs.LG)
  11. Measuring Non-Adversarial Reproduction of Training Data in Large Language Models
    2024/11/15 by Michael Aerni, Javier Rando, Aerni, Michael +9 · 2 voices · 2 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.LG
  12. Gradient-based Jailbreak Images for Multimodal Fusion Models
    2024/10/04 by Javier Rando, Rando, Javier, Hannah Korevaar +7 · 2 citations
    Engineering · Medicine · #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Medical Imaging Techniques and Applications #Medical Imaging and Analysis
  13. Exploring Adversarial Attacks and Defenses in Vision Transformers trained with DINO
    2022/06/14 by Javier Rando, Rando, Javier, Nasib Naimi +5 · 1 citation
    Computer Science · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  14. AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
    2025/03/03 by Nicholas Carlini, Carlini, Nicholas, Javier Rando +7 · 3 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Security and Verification in Computing