vix.ing · top · new · best · stats · spec

Andy Zou

  1. Universal and Transferable Adversarial Attacks on Aligned Language Models
    2023/07/27 by Andy Zou, Zou, Andy, Zifan Wang +9 · 5 voices · 457 citations
    #cs.CL #cs.AI #cs.CR #cs.LG
  2. A Definition of AGI
    2025/10/21 by Dan Hendrycks, Hendrycks, Dan, Dawn Song +65 · 28 voices · 10 citations
    Psychology · Computer Science · #Cognitive Abilities and Testing #Computability, Logic, AI Algorithms #Cognitive Computing and Networks
  3. Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
    2025/12/10 by Justin W. Lin, Eliot Krzysztof Jones, Lin, Justin W. +25 · 24 voices · 4 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Information and Cyber Security #Web Application Security Vulnerabilities #cs.AI #cs.CR #cs.CY
  4. Measuring Massive Multitask Language Understanding
    2020/09/07 by Dan Hendrycks, Collin Burns, Hendrycks, Dan +11 · 1176 citations
    Computer Science · #Topic Modeling #Explainable Artificial Intelligence (XAI) #Natural Language Processing Techniques
  5. Representation Engineering: A Top-Down Approach to AI Transparency
    2023/10/02 by Andy Zou, Long Phan, Zou, Andy +40 · 5 voices · 153 citations
    Computer Science · Engineering · #cs.LG #cs.AI #cs.CL #cs.CV #cs.CY
  6. Lessons from the Trenches on Reproducible Evaluation of Language Models
    2024/05/23 by Stella Biderman, Biderman, Stella, Hailey Schoelkopf +58 · 2 voices · 34 citations
    Computer Science · #Natural Language Processing Techniques
  7. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
    2022/06/09 by Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao +448 · 3 voices · 126 citations
    #cs.CL #cs.AI #cs.CY #cs.LG #stat.ML
  8. Humanity's Last Exam
    2025/01/24 by Long Phan, Phan, Long, Alice Gatti +2240 · 9 voices · 103 citations
    #cs.LG #cs.AI #cs.CL
  9. HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
    2024/02/06 by Mantas Mazeika, Long Phan, Mazeika, Mantas +21 · 230 citations
    Decision Sciences · Computer Science · #Complex Systems and Decision Making #Information and Cyber Security
  10. AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
    2024/10/11 by Maksym Andriushchenko, Alexandra Souly, Andriushchenko, Maksym +25 · 54 citations
    Computer Science · #Blockchain Technology Applications and Security
  11. Tamper-Resistant Safeguards for Open-Weight LLMs
    2024/08/01 by Rishub Tamirisa, Bhrugu Bharathi, Tamirisa, Rishub +27 · 25 citations
    Engineering · #Radiation Effects in Electronics #Electrostatic Discharge in Electronics #Electrical Fault Detection and Protection
  12. Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
    2023/04/06 by Alexander Pan, Pan, Alexander, Andy Zou +16 · 14 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI)
  13. Forecasting Future World Events with Neural Networks
    2022/06/30 by Andy Zou, Zou, Andy, Tristan Xiao +17 · 9 citations
    Computer Science · Social Sciences · #Computation and Language (cs.CL) #Computational and Text Analysis Methods #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
  14. What Would Jiminy Cricket Do? Towards Agents That Behave Morally
    2021/10/25 by Dan Hendrycks, Hendrycks, Dan, Mantas Mazeika +15 · 6 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Artificial Intelligence in Games #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
  15. Safety Pretraining: Toward the Next Generation of Safe AI
    2025/04/23 by Pratyush Maini, Maini, Pratyush, Sachin Goyal +18 · 3 voices · 8 citations
    Health Professions · #cs.LG
  16. Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
    2025/07/28 by Andy Zou, Maxwell Lin, Zou, Andy +31 · 5 voices · 8 citations
    #cs.AI #cs.CL #cs.CY
  17. Transferable Adversarial Attacks on Black-Box Vision-Language Models
    2025/05/02 by Kai Hu, Hu, Kai, Weichen Yu +13 · 3 citations
    Computer Science · #Adversarial Robustness in Machine Learning
  18. TextQuests: How Good are LLMs at Text-Based Video Games?
    2025/07/31 by Long Phan, Phan, Long, Mantas Mazeika +5 · 1 voice · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.AI #cs.CL
  19. D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
    2025/09/22 by Satyapriya Krishna, Krishna, Satyapriya, Andy Zou +13 · 3 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Software Engineering Research #Topic Modeling