vix.ing · top · new · best · stats · spec

Alexander Pan

  1. A Definition of AGI
    2025/10/21 by Dan Hendrycks, Dawn Song, Hendrycks, Dan +65 · 28 voices · 10 citations
    Psychology · Computer Science · #Cognitive Abilities and Testing #Computability, Logic, AI Algorithms #Cognitive Computing and Networks
  2. Representation Engineering: A Top-Down Approach to AI Transparency
    2023/10/02 by Andy Zou, Long Phan, Zou, Andy +40 · 5 voices · 177 citations
    Computer Science · Engineering · #cs.LG #cs.AI #cs.CL #cs.CV #cs.CY
  3. The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
    2024/03/05 by Nathaniel Li, Alexander Pan, Li, Nathaniel +106 · 85 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Network Security and Intrusion Detection
  4. The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
    2022/01/10 by Alexander Pan, Kush Bhatia, Pan, Alexander +3 · 39 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Information and Cyber Security #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Network Security and Intrusion Detection #Reinforcement Learning in Robotics
  5. Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
    2024/07/31 by Richard Ren, Steven Basart, Ren, Richard +21 · 2 voices · 18 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.CY #cs.LG
  6. Feedback Loops With Language Models Drive In-Context Reward Hacking
    2024/02/09 by Alexander Pan, Erik Jones, Pan, Alexander +5 · 1 voice · 17 citations
    Computer Science · #Advanced Malware Detection Techniques #Software Testing and Debugging Techniques #Formal Methods in Verification
  7. Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
    2023/04/06 by Alexander Pan, Pan, Alexander, Andy Zou +16 · 15 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI)
  8. LatentQA: Teaching LLMs to Decode Activations Into Natural Language
    2024/12/11 by Alexander Pan, Lijie Chen, Pan, Alexander +3 · 9 citations
    Computer Science · Social Sciences · #Artificial Intelligence in Law #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling