Kai Fronsdal
- Evaluating whether AI models would sabotage AI safety research
2026/04/27 by Robert Kirk, Alexandra Souly, Kai Fronsdal +2 · 1 voice
Computer Science · Psychology · Social Sciences · #Adversarial Robustness in Machine Learning #CLARITY #Continuation #Covert #Ethics and Social Impacts of AI #Flagging #Frontier #Human-Automation Interaction and Safety #Opus #Similarity (geometry) #Situational ethics #cs.AI