Sebastian Farquhar
- The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation
2018/02/20 by Miles Brundage, Brundage, Miles, Shahar Avin +49 · 6 voices · 19 citations
#cs.AI #cs.CR #cs.CY
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
2023/02/19 by Lorenz Kuhn, Yarin Gal, Kuhn, Lorenz +3 · 99 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Explainable Artificial Intelligence (XAI)
- Evaluating Frontier Models for Dangerous Capabilities
2024/03/20 by Mary Phuong, Phuong, Mary, Matthew Aitchison +51 · 1 voice · 21 citations
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.LG
- Do Bayesian Neural Networks Need To Be Fully Stochastic?
2022/11/11 by Mrinank Sharma, Sharma, Mrinank, Sebastian Farquhar +5 · 1 voice · 7 citations
Computer Science · Mathematics · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #cs.AI #cs.LG #stat.ML
- Tracr: Compiled Transformers as a Laboratory for Interpretability
2023/01/12 by David Lindner, Lindner, David, János Kramár +9 · 1 voice · 5 citations
Computer Science · Mathematics · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification #cs.AI #cs.LG #stat.ML
- Active Testing: Sample-Efficient Model Evaluation
2021/03/09 by Jannik Kossen, Kossen, Jannik, Sebastian Farquhar +5 · 10 citations
Computer Science · #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Machine Learning and Data Classification
- Model evaluation for extreme risks
2023/05/24 by Toby Shevlane, Shevlane, Toby, Sebastian Farquhar +39 · 14 citations
Computer Science · #Software Engineering Research #Software Reliability and Analysis Research #Information and Cyber Security
- Uncertainty Baselines: Benchmarks for Uncertainty & Robustness in Deep Learning
2021/06/07 by Zachary Nado, Nado, Zachary, Neil Band +49 · 5 citations
Computer Science · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification
- A Systematic Comparison of Bayesian Deep Learning Robustness in Diabetic Retinopathy Tasks
2019/12/22 by Angelos Filos, Filos, Angelos, Sebastian Farquhar +15 · 7 citations
Computer Science · Medicine · #Anomaly Detection Techniques and Applications #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Video Processing (eess.IV) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification #Retinal Imaging and Analysis #electronic engineering #information engineering
- Challenges with unsupervised LLM knowledge discovery
2023/12/15 by Sebastian Farquhar, Farquhar, Sebastian, Vikrant Varma +9 · 6 citations
Computer Science · Materials Science · #Topic Modeling #Natural Language Processing Techniques #Machine Learning in Materials Science
- Holistic Safety and Responsibility Evaluations of Advanced AI Models
2024/04/22 by Laura Weidinger, Weidinger, Laura, Joslyn Barnhart +35 · 6 citations
Computer Science · Decision Sciences · #Adversarial Robustness in Machine Learning #Risk and Safety Analysis
- MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
2025/01/22 by Sebastian Farquhar, Vikrant Varma, Farquhar, Sebastian +11 · 1 voice · 4 citations
Computer Science · #Blockchain Technology Applications and Security #cs.AI #cs.LG
- Liberty or Depth: Deep Bayesian Neural Nets Do Not Need Complex Weight Posterior Approximations
2020/02/10 by Sebastian Farquhar, Farquhar, Sebastian, Lewis Smith +3 · 2 citations
Computer Science · Mathematics · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Markov Chains and Monte Carlo Methods
- Path-Specific Objectives for Safer Agent Incentives
2022/04/21 by Sebastian Farquhar, Ryan M. Carey, Farquhar, Sebastian +3 · 1 citation
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Bayesian Modeling and Causal Inference #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Machine Learning (stat.ML)