Skalse, Joar
- Risks from Learned Optimization in Advanced Machine Learning Systems
2019/06/05 by Evan Hubinger, Hubinger, Evan, Chris van Merwijk +7 · 38 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning and Algorithms #Reinforcement Learning in Robotics
- Defining and Characterizing Reward Hacking
2022/09/27 by Joar Skalse, Nikolaus H. R. Howe, Skalse, Joar +5 · 30 citations
Computer Science · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Formal Methods in Verification #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Software Reliability and Analysis Research
- Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems
2024/05/10 by David "davidad" Dalrymple, Joar Skalse, Dalrymple, David "davidad" +31 · 3 voices · 11 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #cs.AI
- Invariance in Policy Optimisation and Partial Identifiability in Reward Learning
2022/03/14 by Joar Skalse, Skalse, Joar, Matthew Farrugia-Roberts +7 · 5 citations
Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Data Quality and Management #Data Stream Mining Techniques #FOS: Computer and information sciences #I.2.6 #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Time Series Analysis and Forecasting
- Goodhart's Law in Reinforcement Learning
2023/10/13 by Jacek Karwowski, Oliver Hayman, Karwowski, Jacek +9 · 3 citations
Business, Management and Accounting · Decision Sciences · #Auction Theory and Applications #FOS: Computer and information sciences #Game Theory and Applications #Machine Learning (cs.LG) #Supply Chain and Inventory Management
- Neural networks are a priori biased towards Boolean functions with low entropy
2019/09/25 by Mingard, Chris, Skalse, Joar, Valle-Pérez, Guillermo +3 · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Partial Identifiability and Misspecification in Inverse Reinforcement Learning
2024/11/24 by Joar Skalse, Alessandro Abate, Skalse, Joar +1 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics
- STARC: A General Framework For Quantifying Differences Between Reward Functions
2023/09/26 by Joar Skalse, Skalse, Joar, Lucy Farnik +9 · 1 citation
Biochemistry, Genetics and Molecular Biology · Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Receptor Mechanisms and Signaling #Reinforcement Learning in Robotics
- The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret
2024/06/22 by Lukas Fluri, Leon Lang, Fluri, Lukas +9 · 1 voice · 1 citation
Psychology · #Human Resource Development and Performance Evaluation #cs.AI #cs.LG #stat.ML