Dan Hendrycks
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
2025/07/15 by Tomek Korbak, Mikita Balesni, Korbak, Tomek +79 · 25 voices · 66 citations
#cs.AI #cs.LG #stat.ML
- Unsolved Problems in ML Safety
2021/09/28 by Dan Hendrycks, Nicholas Carlini, Hendrycks, Dan +5 · 2 voices · 27 citations
#cs.LG #cs.AI #cs.CL #cs.CV
- A Definition of AGI
2025/10/21 by Dan Hendrycks, Hendrycks, Dan, Dawn Song +65 · 28 voices · 10 citations
Psychology · Computer Science · #Cognitive Abilities and Testing #Computability, Logic, AI Algorithms #Cognitive Computing and Networks
- Measuring Massive Multitask Language Understanding
2020/09/07 by Dan Hendrycks, Hendrycks, Dan, Collin Burns +11 · 1174 citations
Computer Science · #Topic Modeling #Explainable Artificial Intelligence (XAI) #Natural Language Processing Techniques
- An Overview of Catastrophic AI Risks
2023/06/21 by Dan Hendrycks, Mantas Mazeika, Hendrycks, Dan +3 · 5 voices · 25 citations
#cs.CY #cs.AI #cs.LG
- Measuring Mathematical Problem Solving With the MATH Dataset
2021/03/05 by Dan Hendrycks, Hendrycks, Dan, Collin Burns +13 · 1030 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Advanced Text Analysis Techniques
- Representation Engineering: A Top-Down Approach to AI Transparency
2023/10/02 by Andy Zou, Long Phan, Zou, Andy +40 · 5 voices · 153 citations
Computer Science · Engineering · #cs.LG #cs.AI #cs.CL #cs.CV #cs.CY
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
2022/06/09 by Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao +448 · 3 voices · 125 citations
#cs.CL #cs.AI #cs.CY #cs.LG #stat.ML
- Gaussian Error Linear Units (GELUs)
2016/06/27 by Dan Hendrycks, Hendrycks, Dan, Kevin Gimpel +1 · 318 citations
Computer Science · #Advanced Neural Network Applications #Anomaly Detection Techniques and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
2025/02/12 by Mantas Mazeika, Xuwang Yin, Mazeika, Mantas +21 · 18 voices · 17 citations
Computer Science · Engineering · Decision Sciences · #AI-based Problem Solving and Planning #Flexible and Reconfigurable Manufacturing Systems #Simulation Techniques and Applications
- Natural Selection Favors AIs over Humans
2023/03/28 by Dan Hendrycks, Hendrycks, Dan · 8 voices · 1 citation
#cs.CY #cs.AI #cs.LG #cs.NE
- Benchmarking Neural Network Robustness to Common Corruptions and Perturbations
2019/03/28 by Dan Hendrycks, Thomas G. Dietterich, Hendrycks, Dan +1 · 209 citations
Computer Science · #Adversarial Robustness in Machine Learning #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization
2020/06/29 by Dan Hendrycks, Steven Basart, Hendrycks, Dan +23 · 153 citations
Computer Science · Biochemistry, Genetics and Molecular Biology · #Advanced Image Processing Techniques #Image and Signal Denoising Methods #Cell Image Analysis Techniques
- Humanity's Last Exam
2025/01/24 by Long Phan, Phan, Long, Alice Gatti +2240 · 9 voices · 103 citations
#cs.LG #cs.AI #cs.CL
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
2024/02/06 by Mantas Mazeika, Long Phan, Mazeika, Mantas +21 · 230 citations
Decision Sciences · Computer Science · #Complex Systems and Decision Making #Information and Cyber Security
- Natural Adversarial Examples
2019/07/16 by Dan Hendrycks, Hendrycks, Dan, Kevin Zhao +7 · 118 citations
Computer Science · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Aligning AI With Shared Human Values
2020/08/05 by Dan Hendrycks, Collin Burns, Hendrycks, Dan +11 · 92 citations
Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- Deep Anomaly Detection with Outlier Exposure
2018/12/11 by Dan Hendrycks, Mantas Mazeika, Hendrycks, Dan +3 · 97 citations
Computer Science · #Advanced Malware Detection Techniques #Anomaly Detection Techniques and Applications #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Network Security and Intrusion Detection
- DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
2023/06/20 by Boxin Wang, Weixin Chen, Wang, Boxin +35 · 58 citations
Computer Science · Medicine · Social Sciences · #Adversarial Robustness in Machine Learning #Artificial Intelligence in Healthcare and Education #Ethics and Social Impacts of AI
- AI Deception: A Survey of Examples, Risks, and Potential Solutions
2023/08/28 by Peter S. Park, Park, Peter S., Simon Goldstein +7 · 36 citations
Social Sciences · Computer Science · #Ethics and Social Impacts of AI #Cybercrime and Law Enforcement Studies #Crime, Illicit Activities, and Governance
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
2024/10/11 by Maksym Andriushchenko, Andriushchenko, Maksym, Alexandra Souly +25 · 54 citations
Computer Science · #Blockchain Technology Applications and Security
- Using Self-Supervised Learning Can Improve Model Robustness and Uncertainty
2019/06/28 by Dan Hendrycks, Hendrycks, Dan, Mantas Mazeika +5 · 20 citations
Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification
- Remote Labor Index: Measuring AI Automation of Remote Work
2025/10/30 by Mantas Mazeika, Alice Gatti, Mazeika, Mantas +91 · 9 voices · 9 citations
#cs.LG #cs.AI #cs.CL
- Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
2024/07/31 by Richard Ren, Ren, Richard, Steven Basart +21 · 2 voices · 14 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.CY #cs.LG
- Using Trusted Data to Train Deep Networks on Labels Corrupted by Severe Noise
2018/02/14 by Dan Hendrycks, Hendrycks, Dan, Mantas Mazeika +5 · 21 citations
Computer Science · #Machine Learning and Data Classification #Adversarial Robustness in Machine Learning #Machine Learning and Algorithms
- Benchmarking Neural Network Robustness to Common Corruptions and Surface Variations
2018/07/04 by Dan Hendrycks, Hendrycks, Dan, Thomas G. Dietterich +1 · 16 citations
Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural and Evolutionary Computing (cs.NE)
- Tamper-Resistant Safeguards for Open-Weight LLMs
2024/08/01 by Rishub Tamirisa, Tamirisa, Rishub, Bhrugu Bharathi +27 · 25 citations
Engineering · #Radiation Effects in Electronics #Electrostatic Discharge in Electronics #Electrical Fault Detection and Protection
- EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges
2025/02/13 by Clinton J. Wang, Dean Lee, Wang, Clinton J. +18 · 3 voices · 4 citations
Computer Science · #Natural Language Processing Techniques
- Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
2023/04/06 by Alexander Pan, Pan, Alexander, Chan, Jun Shern +16 · 14 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI)
- A Unified Survey on Anomaly, Novelty, Open-Set, and Out-of-Distribution Detection: Solutions and Future Challenges
2021/10/26 by Mohammadreza Salehi, Hossein Mirzaei, Salehi, Mohammadreza +9 · 10 citations
Computer Science · Medicine · #Anomaly Detection Techniques and Applications #Data-Driven Disease Surveillance #Machine Learning and Data Classification
- Forecasting Future World Events with Neural Networks
2022/06/30 by Andy Zou, Zou, Andy, Tristan Xiao +17 · 9 citations
Computer Science · Social Sciences · #Computation and Language (cs.CL) #Computational and Text Analysis Methods #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- What Would Jiminy Cricket Do? Towards Agents That Behave Morally
2021/10/25 by Dan Hendrycks, Hendrycks, Dan, Mantas Mazeika +15 · 6 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Artificial Intelligence in Games #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
2024/03/18 by Junyuan Hong, Jinhao Duan, Hong, Junyuan +27 · 9 citations
Computer Science · #Advanced Data Storage Technologies #Artificial Intelligence (cs.AI) #Blockchain Technology Applications and Security #Computation and Language (cs.CL) #Cryptography and Data Security #FOS: Computer and information sciences
- LLM-PBE: Assessing Data Privacy in Large Language Models
2024/08/23 by Qinbin Li, Junyuan Hong, Li, Qinbin +23 · 11 citations
Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #Data Quality and Management #FOS: Computer and information sciences #Privacy-Preserving Technologies in Data
- The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
2025/03/05 by Richard Ren, Ren, Richard, Mantas Mazeika +26 · 14 citations
Computer Science · Social Sciences · #Explainable Artificial Intelligence (XAI) #Ethics and Social Impacts of AI #Adversarial Robustness in Machine Learning
- Can LLMs Follow Simple Rules?
2023/11/06 by Norman Mu, Mu, Norman, Sarah Chen +14 · 6 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Software Engineering Research #Topic Modeling
- MAUD: An Expert-Annotated Legal NLP Dataset for Merger Agreement Understanding
2023/01/02 by Steven H. Wang, Antoine Scardigli, Wang, Steven H. +17 · 4 citations
Social Sciences · Computer Science · #Artificial Intelligence in Law #Law, AI, and Intellectual Property #Legal Education and Practice Innovations
- Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
2025/07/28 by Andy Zou, Zou, Andy, Maxwell Lin +31 · 5 voices · 8 citations
#cs.AI #cs.CL #cs.CY
- The Singapore Consensus on Global AI Safety Research Priorities
2025/06/25 by Yoshua Bengio, Tegan Maharaj, Bengio, Yoshua +171 · 2 voices · 8 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences #cs.AI #cs.CY
- Open Category Detection with PAC Guarantees
2018/08/01 by Si Liu, Risheek Garrepalli, Liu, Si +7 · 3 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Anomaly Detection Techniques and Applications #Machine Learning and Algorithms
- Introduction to AI Safety, Ethics, and Society
2024/11/01 by Dan Hendrycks, Hendrycks, Dan · 6 citations
Social Sciences · #Ethics and Social Impacts of AI
- Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark
2025/04/21 by Jasper Götting, Pedro Henrique Quintela Soares de Medeiros, Götting, Jasper +15 · 8 citations
Agricultural and Biological Sciences · Medicine · #Animal Disease Management and Epidemiology #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Respiratory viral infections research
- Early Methods for Detecting Adversarial Images
2016/08/01 by Dan Hendrycks, Kevin Gimpel, Hendrycks, Dan +1 · 1 citation
Biochemistry, Genetics and Molecular Biology · Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Cell Image Analysis Techniques #Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE)
- Beyond Release: Access Considerations for Generative AI Systems
2025/02/23 by Irene Solaiman, Solaiman, Irene, Rishi Bommasani +11 · 3 citations
Computer Science · Decision Sciences · #Semantic Web and Ontologies #Computability, Logic, AI Algorithms #Scientific Computing and Data Management
- Testing Robustness Against Unforeseen Adversaries
2019/08/21 by Max Kaufmann, Kaufmann, Max, Daniel Kang +21 · 1 voice
Computer Science · Mathematics · #Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #cs.CR #cs.CV #cs.LG #stat.ML
- MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models
2025/03/19 by Chejian Xu, Jiawei Zhang, Xu, Chejian +44 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Multi-Agent Systems and Negotiation
- TextQuests: How Good are LLMs at Text-Based Video Games?
2025/07/31 by Long Phan, Phan, Long, Mantas Mazeika +5 · 1 voice · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.AI #cs.CL
- D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models
2025/09/22 by Satyapriya Krishna, Andy Zou, Krishna, Satyapriya +13 · 3 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Software Engineering Research #Topic Modeling