Alexander Pan
- A Definition of AGI
2025/10/21 by Dan Hendrycks, Dawn Song, Hendrycks, Dan +65 · 28 voices · 10 citations
Psychology · Computer Science · #Cognitive Abilities and Testing #Computability, Logic, AI Algorithms #Cognitive Computing and Networks
- Representation Engineering: A Top-Down Approach to AI Transparency
2023/10/02 by Andy Zou, Long Phan, Zou, Andy +40 · 5 voices · 177 citations
Computer Science · Engineering · #cs.LG #cs.AI #cs.CL #cs.CV #cs.CY
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
2024/03/05 by Nathaniel Li, Alexander Pan, Li, Nathaniel +106 · 85 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Network Security and Intrusion Detection
- The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
2022/01/10 by Alexander Pan, Kush Bhatia, Pan, Alexander +3 · 39 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Information and Cyber Security #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Network Security and Intrusion Detection #Reinforcement Learning in Robotics
- Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
2024/07/31 by Richard Ren, Steven Basart, Ren, Richard +21 · 2 voices · 18 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.CY #cs.LG
- Feedback Loops With Language Models Drive In-Context Reward Hacking
2024/02/09 by Alexander Pan, Erik Jones, Pan, Alexander +5 · 1 voice · 17 citations
Computer Science · #Advanced Malware Detection Techniques #Software Testing and Debugging Techniques #Formal Methods in Verification
- Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
2023/04/06 by Alexander Pan, Pan, Alexander, Andy Zou +16 · 15 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI)
- LatentQA: Teaching LLMs to Decode Activations Into Natural Language
2024/12/11 by Alexander Pan, Lijie Chen, Pan, Alexander +3 · 9 citations
Computer Science · Social Sciences · #Artificial Intelligence in Law #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling