vix.ing · top · new · best · stats · spec

Ren, Richard

  1. Representation Engineering: A Top-Down Approach to AI Transparency
    2023/10/02 by Andy Zou, Long Phan, Zou, Andy +40 · 5 voices · 193 citations
    Computer Science · Engineering · #cs.LG #cs.AI #cs.CL #cs.CV #cs.CY
  2. Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
    2025/02/12 by Mantas Mazeika, Mazeika, Mantas, Xuwang Yin +21 · 18 voices · 20 citations
    Computer Science · Engineering · Decision Sciences · #AI-based Problem Solving and Planning #Flexible and Reconfigurable Manufacturing Systems #Simulation Techniques and Applications
  3. Humanity's Last Exam
    2025/01/24 by Long Phan, Alice Gatti, Phan, Long +2240 · 9 voices · 120 citations
    #cs.LG #cs.AI #cs.CL
  4. Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
    2024/07/31 by Richard Ren, Steven Basart, Ren, Richard +22 · 2 voices · 20 citations
    Computer Science · Health Professions · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Occupational Health and Safety Research #cs.AI #cs.CL #cs.CY #cs.LG
  5. Remote Labor Index: Measuring AI Automation of Remote Work
    2025/10/30 by Mantas Mazeika, Alice Gatti, Mazeika, Mantas +91 · 9 voices · 9 citations
    #cs.LG #cs.AI #cs.CL
  6. The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems
    2025/03/05 by Richard Ren, Ren, Richard, Mantas Mazeika +26 · 17 citations
    Computer Science · Social Sciences · #Explainable Artificial Intelligence (XAI) #Ethics and Social Impacts of AI #Adversarial Robustness in Machine Learning
  7. Localizing Lying in Llama: Understanding Instructed Dishonesty on True-False Questions Through Prompting, Probing, and Patching
    2023/11/25 by James Campbell, Richard Ren, Campbell, James +3 · 4 citations
    Computer Science · Psychology · Social Sciences · #Topic Modeling #Deception detection and forensic psychology #Ethics and Social Impacts of AI
  8. Deep Reinforcement Learning for Decentralized Multi-Robot Exploration With Macro Actions
    2021/10/05 by Aaron Hao Tan, Tan, Aaron Hao, Federico Pizarro Bejarano +4 · 3 citations
    Computer Science · Psychology · Social Sciences · #Evolutionary Game Theory and Cooperation #FOS: Computer and information sciences #Reinforcement Learning in Robotics #Robotics (cs.RO) #Social Robot Interaction and HRI