vix.ing · top · new · best · stats · spec

Skalse, Joar

  1. Risks from Learned Optimization in Advanced Machine Learning Systems
    2019/06/05 by Evan Hubinger, Hubinger, Evan, Chris van Merwijk +7 · 38 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning and Algorithms #Reinforcement Learning in Robotics
  2. Defining and Characterizing Reward Hacking
    2022/09/27 by Joar Skalse, Nikolaus H. R. Howe, Skalse, Joar +5 · 30 citations
    Computer Science · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Formal Methods in Verification #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Software Reliability and Analysis Research
  3. Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems
    2024/05/10 by David "davidad" Dalrymple, Joar Skalse, Dalrymple, David "davidad" +31 · 3 voices · 11 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #cs.AI
  4. Invariance in Policy Optimisation and Partial Identifiability in Reward Learning
    2022/03/14 by Joar Skalse, Skalse, Joar, Matthew Farrugia-Roberts +7 · 5 citations
    Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Data Quality and Management #Data Stream Mining Techniques #FOS: Computer and information sciences #I.2.6 #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Time Series Analysis and Forecasting
  5. Goodhart's Law in Reinforcement Learning
    2023/10/13 by Jacek Karwowski, Oliver Hayman, Karwowski, Jacek +9 · 3 citations
    Business, Management and Accounting · Decision Sciences · #Auction Theory and Applications #FOS: Computer and information sciences #Game Theory and Applications #Machine Learning (cs.LG) #Supply Chain and Inventory Management
  6. Neural networks are a priori biased towards Boolean functions with low entropy
    2019/09/25 by Mingard, Chris, Skalse, Joar, Valle-Pérez, Guillermo +3 · 1 citation
    #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  7. Partial Identifiability and Misspecification in Inverse Reinforcement Learning
    2024/11/24 by Joar Skalse, Alessandro Abate, Skalse, Joar +1 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics
  8. STARC: A General Framework For Quantifying Differences Between Reward Functions
    2023/09/26 by Joar Skalse, Skalse, Joar, Lucy Farnik +9 · 1 citation
    Biochemistry, Genetics and Molecular Biology · Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Receptor Mechanisms and Signaling #Reinforcement Learning in Robotics
  9. The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret
    2024/06/22 by Lukas Fluri, Leon Lang, Fluri, Lukas +9 · 1 voice · 1 citation
    Psychology · #Human Resource Development and Performance Evaluation #cs.AI #cs.LG #stat.ML