vix.ing · top · new · best · stats · spec

Charbel-Raphaël Segerie

  1. Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
    2023/07/27 by Stephen Casper, Casper, Stephen, Xander Davies +65 · 3 voices · 147 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Reliability and Analysis Research #cs.AI #cs.CL #cs.LG
  2. Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods
    2025/05/08 by Charbel-Raphaël Segerie, Grey, Markov, Segerie, Charbel-Raphaël · 10 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI)
  3. BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM Safeguards
    2024/06/03 by Diego Dorn, Alexandre Variengien, Dorn, Diego +5 · 3 citations
    Engineering · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Nuclear and radioactivity studies
  4. In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?
    2025/04/17 by Ben Bucknall, Saad Siddiqui, Bucknall, Ben +41 · 2 voices · 2 citations
    #cs.CY