Korbak, Tomasz
- AI Supported Degradation of the Self Concept: A Theoretical Framework Grounded in Established Cognitive and Computational Mechanisms
2023/10/20 by Mrinank Sharma, Sharma, Mrinank, Meg Tong +36 · 10 voices · 228 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Reinforcement Learning in Robotics #Topic Modeling #cs.AI #cs.CL #cs.LG #stat.ML
- Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
2024/04/01 by Matthias Gerstgrasser, Rylan Schaeffer, Gerstgrasser, Matthias +26 · 16 voices · 31 citations
Computer Science · #Semantic Web and Ontologies #cs.AI #cs.CL #cs.ET #cs.LG #stat.ML
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
2023/09/21 by Lukas Berglund, Berglund, Lukas, Meg Tong +11 · 11 voices · 76 citations
Computer Science · #Law, AI, and Intellectual Property #cs.AI #cs.CL #cs.LG
- Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
2023/07/27 by Stephen Casper, Xander Davies, Casper, Stephen +65 · 3 voices · 89 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Reliability and Analysis Research #cs.AI #cs.CL #cs.LG
- Pretraining Language Models with Human Preferences
2023/02/16 by Tomasz Korbak, Korbak, Tomasz, Kejian Shi +13 · 1 voice · 17 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.LG
- RL with KL penalties is better viewed as Bayesian inference
2022/05/23 by Tomasz Korbak, Ethan Perez, Korbak, Tomasz +4 · 2 voices · 18 citations
Computer Science · Mathematics · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Topic Modeling #cs.LG #stat.ML
- Aligning Language Models with Preferences through f-divergence Minimization
2023/02/16 by Go, Dongyoung, Korbak, Tomasz, Kruszewski, Germán +3 · 19 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Foundational Challenges in Assuring Alignment and Safety of Large Language Models
2024/04/15 by Anwar, Usman, Saparov, Abulhair, Rando, Javier +39 · 22 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Taken out of context: On measuring situational awareness in LLMs
2023/09/01 by Lukas Berglund, Berglund, Lukas, Asa Cooper Stickland +13 · 13 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
- On Reinforcement Learning and Distribution Matching for Fine-Tuning Language Models with no Catastrophic Forgetting
2022/06/01 by Tomasz Korbak, Korbak, Tomasz, Hady Elsahar +5 · 9 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- Improving Code Generation by Training with Natural Language Feedback
2023/03/28 by Chen, Angelica, Scheurer, Jérémy, Korbak, Tomasz +5 · 4 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Engineering (cs.SE)
- A continuity of Markov blanket interpretations under the Free Energy Principle
2022/01/18 by Seth, Anil, Korbak, Tomasz, Tschantz, Alexander · 3 citations
#FOS: Biological sciences #Neurons and Cognition (q-bio.NC)
- Training Language Models with Language Feedback at Scale
2023/03/28 by Jérémy Scheurer, Scheurer, Jérémy, Jon Ander Campos +11 · 4 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Explainable Artificial Intelligence (XAI)
- Controlling Conditional Language Models without Catastrophic Forgetting
2021/12/01 by Tomasz Korbak, Korbak, Tomasz, Hady Elsahar +5 · 3 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- Measuring non-trivial compositionality in emergent communication
2020/10/28 by Korbak, Tomasz, Zubek, Julian, Rączaszek-Leonardi, Joanna · 2 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE)