vix.ing · top · new · best · stats · spec

Korbak, Tomasz

  1. AI Supported Degradation of the Self Concept: A Theoretical Framework Grounded in Established Cognitive and Computational Mechanisms
    2023/10/20 by Mrinank Sharma, Sharma, Mrinank, Meg Tong +36 · 10 voices · 228 citations
    Computer Science · #Explainable Artificial Intelligence (XAI) #Reinforcement Learning in Robotics #Topic Modeling #cs.AI #cs.CL #cs.LG #stat.ML
  2. Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
    2024/04/01 by Matthias Gerstgrasser, Rylan Schaeffer, Gerstgrasser, Matthias +26 · 16 voices · 31 citations
    Computer Science · #Semantic Web and Ontologies #cs.AI #cs.CL #cs.ET #cs.LG #stat.ML
  3. The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
    2023/09/21 by Lukas Berglund, Berglund, Lukas, Meg Tong +11 · 11 voices · 76 citations
    Computer Science · #Law, AI, and Intellectual Property #cs.AI #cs.CL #cs.LG
  4. Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
    2023/07/27 by Stephen Casper, Xander Davies, Casper, Stephen +65 · 3 voices · 89 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Reliability and Analysis Research #cs.AI #cs.CL #cs.LG
  5. Pretraining Language Models with Human Preferences
    2023/02/16 by Tomasz Korbak, Korbak, Tomasz, Kejian Shi +13 · 1 voice · 17 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.LG
  6. RL with KL penalties is better viewed as Bayesian inference
    2022/05/23 by Tomasz Korbak, Ethan Perez, Korbak, Tomasz +4 · 2 voices · 18 citations
    Computer Science · Mathematics · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Topic Modeling #cs.LG #stat.ML
  7. Aligning Language Models with Preferences through f-divergence Minimization
    2023/02/16 by Go, Dongyoung, Korbak, Tomasz, Kruszewski, Germán +3 · 19 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  8. Foundational Challenges in Assuring Alignment and Safety of Large Language Models
    2024/04/15 by Anwar, Usman, Saparov, Abulhair, Rando, Javier +39 · 22 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  9. Taken out of context: On measuring situational awareness in LLMs
    2023/09/01 by Lukas Berglund, Berglund, Lukas, Asa Cooper Stickland +13 · 13 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
  10. On Reinforcement Learning and Distribution Matching for Fine-Tuning Language Models with no Catastrophic Forgetting
    2022/06/01 by Tomasz Korbak, Korbak, Tomasz, Hady Elsahar +5 · 9 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  11. Improving Code Generation by Training with Natural Language Feedback
    2023/03/28 by Chen, Angelica, Scheurer, Jérémy, Korbak, Tomasz +5 · 4 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Engineering (cs.SE)
  12. A continuity of Markov blanket interpretations under the Free Energy Principle
    2022/01/18 by Seth, Anil, Korbak, Tomasz, Tschantz, Alexander · 3 citations
    #FOS: Biological sciences #Neurons and Cognition (q-bio.NC)
  13. Training Language Models with Language Feedback at Scale
    2023/03/28 by Jérémy Scheurer, Scheurer, Jérémy, Jon Ander Campos +11 · 4 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Explainable Artificial Intelligence (XAI)
  14. Controlling Conditional Language Models without Catastrophic Forgetting
    2021/12/01 by Tomasz Korbak, Korbak, Tomasz, Hady Elsahar +5 · 3 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
  15. Measuring non-trivial compositionality in emergent communication
    2020/10/28 by Korbak, Tomasz, Zubek, Julian, Rączaszek-Leonardi, Joanna · 2 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE)