Robert Kirk
- Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models
2024/11/19 by Laura Ruis, Maximilian Mozes, Ruis, Laura +18 · 24 voices · 8 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Semantic Web and Ontologies
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
2025/10/08 by Alexandra Souly, Javier Rando, Souly, Alexandra +25 · 34 voices · 19 citations
Medicine · #Medical Imaging and Pathology Studies
- Understanding the Effects of RLHF on LLM Generalisation and Diversity
2023/10/10 by Robert Kirk, Kirk, Robert, Ishita Mediratta +12 · 1 voice · 74 citations
Computer Science · #cs.LG #cs.AI #cs.CL
- Open Problems in Machine Unlearning for AI Safety
2025/01/09 by Fazl Barez, Tingchen Fu, Barez, Fazl +39 · 6 voices · 11 citations
Engineering · #Fault Detection and Control Systems
- Reward Model Ensembles Help Mitigate Overoptimization
2023/10/04 by Thomas Coste, Coste, Thomas, Robert Kirk +4 · 25 citations
Computer Science · #Machine Learning and Data Classification #Topic Modeling
- Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
2023/11/21 by Samyak Jain, Jain, Samyak, Robert Kirk +13 · 10 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Topic Modeling #Adversarial Robustness in Machine Learning
- Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
2025/07/28 by Andy Zou, Maxwell Lin, Zou, Andy +31 · 5 voices · 8 citations
#cs.AI #cs.CL #cs.CY
- Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
2025/02/03 by Zora Che, Stephen Casper, Che, Zora +27 · 10 citations
Computer Science · #Security and Verification in Computing #Advanced Malware Detection Techniques #Network Security and Intrusion Detection
- Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
2025/08/08 by Kyle O'Brien, Kyle O’Brien, Stephen Casper +18 · 2 voices · 14 citations
Computer Science · Social Sciences · #cs.LG #cs.AI
- How Do Large Language Monkeys Get Their Power (Laws)?
2025/02/24 by Rylan Schaeffer, Joshua Kazdan, Schaeffer, Rylan +14 · 10 citations
Social Sciences · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Language and cultural evolution #Machine Learning (cs.LG)
- Existing Large Language Model Unlearning Evaluations Are Inconclusive
2025/05/31 by Zhili Feng, Feng, Zhili, Ye Xu +13 · 4 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling
- STACK: Adversarial Attacks on LLM Safeguard Pipelines
2025/06/30 by Ian R. McKenzie, Oskar J. Hollinsworth, McKenzie, Ian R. +13 · 3 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Ethics and Social Impacts of AI
- Dataset Featurization: Uncovering Natural Language Features through Unsupervised Data Reconstruction
2025/02/24 by Michal Bravansky, Bravansky, Michal, Suhas Hariharan +4 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques