Andy Zou
- Universal and Transferable Adversarial Attacks on Aligned Language Models
2023/07/27 by Andy Zou, Zou, Andy, Zifan Wang +9 · 5 voices · 457 citations
#cs.CL #cs.AI #cs.CR #cs.LG
- A Definition of AGI
2025/10/21 by Dan Hendrycks, Hendrycks, Dan, Dawn Song +65 · 28 voices · 10 citations
Psychology · Computer Science · #Cognitive Abilities and Testing #Computability, Logic, AI Algorithms #Cognitive Computing and Networks
- Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
2025/12/10 by Justin W. Lin, Eliot Krzysztof Jones, Lin, Justin W. +25 · 24 voices · 4 citations
Computer Science · #Adversarial Robustness in Machine Learning #Information and Cyber Security #Web Application Security Vulnerabilities #cs.AI #cs.CR #cs.CY
- Measuring Massive Multitask Language Understanding
2020/09/07 by Dan Hendrycks, Collin Burns, Hendrycks, Dan +11 · 1176 citations
Computer Science · #Topic Modeling #Explainable Artificial Intelligence (XAI) #Natural Language Processing Techniques
- Representation Engineering: A Top-Down Approach to AI Transparency
2023/10/02 by Andy Zou, Long Phan, Zou, Andy +40 · 5 voices · 153 citations
Computer Science · Engineering · #cs.LG #cs.AI #cs.CL #cs.CV #cs.CY
- Lessons from the Trenches on Reproducible Evaluation of Language Models
2024/05/23 by Stella Biderman, Biderman, Stella, Hailey Schoelkopf +58 · 2 voices · 34 citations
Computer Science · #Natural Language Processing Techniques
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
2022/06/09 by Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao +448 · 3 voices · 126 citations
#cs.CL #cs.AI #cs.CY #cs.LG #stat.ML
- Humanity's Last Exam
2025/01/24 by Long Phan, Phan, Long, Alice Gatti +2240 · 9 voices · 103 citations
#cs.LG #cs.AI #cs.CL
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
2024/02/06 by Mantas Mazeika, Long Phan, Mazeika, Mantas +21 · 230 citations
Decision Sciences · Computer Science · #Complex Systems and Decision Making #Information and Cyber Security
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
2024/10/11 by Maksym Andriushchenko, Alexandra Souly, Andriushchenko, Maksym +25 · 54 citations
Computer Science · #Blockchain Technology Applications and Security
- Tamper-Resistant Safeguards for Open-Weight LLMs
2024/08/01 by Rishub Tamirisa, Bhrugu Bharathi, Tamirisa, Rishub +27 · 25 citations
Engineering · #Radiation Effects in Electronics #Electrostatic Discharge in Electronics #Electrical Fault Detection and Protection
- Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
2023/04/06 by Alexander Pan, Pan, Alexander, Andy Zou +16 · 14 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI)
- Forecasting Future World Events with Neural Networks
2022/06/30 by Andy Zou, Zou, Andy, Tristan Xiao +17 · 9 citations
Computer Science · Social Sciences · #Computation and Language (cs.CL) #Computational and Text Analysis Methods #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- What Would Jiminy Cricket Do? Towards Agents That Behave Morally
2021/10/25 by Dan Hendrycks, Hendrycks, Dan, Mantas Mazeika +15 · 6 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Artificial Intelligence in Games #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- Safety Pretraining: Toward the Next Generation of Safe AI
2025/04/23 by Pratyush Maini, Maini, Pratyush, Sachin Goyal +18 · 3 voices · 8 citations
Health Professions · #cs.LG
- Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
2025/07/28 by Andy Zou, Maxwell Lin, Zou, Andy +31 · 5 voices · 8 citations
#cs.AI #cs.CL #cs.CY
- Transferable Adversarial Attacks on Black-Box Vision-Language Models
2025/05/02 by Kai Hu, Hu, Kai, Weichen Yu +13 · 3 citations
Computer Science · #Adversarial Robustness in Machine Learning
- TextQuests: How Good are LLMs at Text-Based Video Games?
2025/07/31 by Long Phan, Phan, Long, Mantas Mazeika +5 · 1 voice · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.AI #cs.CL
- D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
2025/09/22 by Satyapriya Krishna, Krishna, Satyapriya, Andy Zou +13 · 3 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Software Engineering Research #Topic Modeling