Kirk, Robert
- Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
2024/11/19 by Laura Ruis, Ruis, Laura, Maximilian Mozes +18 · 24 voices · 8 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Semantic Web and Ontologies
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
2025/10/08 by Alexandra Souly, Souly, Alexandra, Javier Rando +25 · 34 voices · 18 citations
Medicine · #Medical Imaging and Pathology Studies
- Understanding the Effects of RLHF on LLM Generalisation and Diversity
2023/10/10 by Robert Kirk, Ishita Mediratta, Kirk, Robert +12 · 1 voice · 74 citations
Computer Science · #cs.LG #cs.AI #cs.CL
- Reward Model Ensembles Help Mitigate Overoptimization
2023/10/04 by Thomas Coste, Coste, Thomas, Robert Kirk +4 · 25 citations
Computer Science · #Machine Learning and Data Classification #Topic Modeling
- Open Problems in Machine Unlearning for AI Safety
2025/01/09 by Fazl Barez, Tingchen Fu, Barez, Fazl +39 · 6 voices · 9 citations
Engineering · #Fault Detection and Control Systems
- Analyzing the Generalization and Reliability of Steering Vectors
2024/07/17 by Tan, Daniel, Chanin, David, Lynch, Aengus +4 · 26 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- MiniHack the Planet: A Sandbox for Open-Ended Reinforcement Learning Research
2021/09/27 by Samvelyan, Mikayel, Kirk, Robert, Kurin, Vitaly +7 · 6 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
2023/11/21 by Samyak Jain, Robert Kirk, Jain, Samyak +13 · 9 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Topic Modeling #Adversarial Robustness in Machine Learning
- Investigating Non-Transitivity in LLM-as-a-Judge
2025/02/19 by Xu, Yi, Ruis, Laura, Rocktäschel, Tim +1 · 11 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
2025/07/28 by Andy Zou, Zou, Andy, Maxwell Lin +31 · 5 voices · 8 citations
#cs.AI #cs.CL #cs.CY
- Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
2025/02/03 by Zora Che, Stephen Casper, Che, Zora +27 · 10 citations
Computer Science · #Security and Verification in Computing #Advanced Malware Detection Techniques #Network Security and Intrusion Detection
- How Do Large Language Monkeys Get Their Power (Laws)?
2025/02/24 by Rylan Schaeffer, Joshua Kazdan, Schaeffer, Rylan +14 · 10 citations
Social Sciences · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Language and cultural evolution #Machine Learning (cs.LG)
- Generalization to New Sequential Decision Making Tasks with In-Context Learning
2023/12/06 by Raparthy, Sharath Chandra, Hambro, Eric, Kirk, Robert +2 · 4 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
2025/08/08 by Kyle O’Brien, Kyle O'Brien, O'Brien, Kyle +18 · 2 voices · 13 citations
Computer Science · Social Sciences · #cs.LG #cs.AI
- Existing Large Language Model Unlearning Evaluations Are Inconclusive
2025/05/31 by Zhili Feng, Feng, Zhili, Ye Xu +13 · 4 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling
- STACK: Adversarial Attacks on LLM Safeguard Pipelines
2025/06/30 by Ian R. McKenzie, Oskar J. Hollinsworth, McKenzie, Ian R. +13 · 3 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Ethics and Social Impacts of AI
- Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs
2025/02/20 by Davies, Xander, Winsor, Eric, Souly, Alexandra +4 · 2 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- An Example Safety Case for Safeguards Against Misuse
2025/05/23 by Clymer, Joshua, Weinbaum, Jonah, Kirk, Robert +3 · 2 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Dataset Featurization: Uncovering Natural Language Features through Unsupervised Data Reconstruction
2025/02/24 by Michal Bravansky, Bravansky, Michal, Kubon, Vaclav +4 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques