Salle, Jeanne
- Auditing language models for hidden objectives
2025/03/14 by Samuel D. Marks, Samuel Marks, Johannes Treutlein +71 · 1 voice · 29 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CL #cs.LG