Ann-Kathrin Dombrowski
- Representation Engineering: A Top-Down Approach to AI Transparency
2023/10/02 by Andy Zou, Zou, Andy, Long Phan +40 · 5 voices · 211 citations
Computer Science · Engineering · #cs.LG #cs.AI #cs.CL #cs.CV #cs.CY
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
2024/03/05 by Nathaniel Li, Alexander Pan, Li, Nathaniel +106 · 104 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Network Security and Intrusion Detection
- Explanations can be manipulated and geometry is to blame
2019/06/19 by Ann-Kathrin Dombrowski, Dombrowski, Ann-Kathrin, Maximilian Alber +9 · 24 citations
Computer Science · Decision Sciences · Medicine · #Artificial Intelligence in Healthcare and Education #Cryptography and Security (cs.CR) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Scientific Computing and Data Management
- Fairwashing Explanations with Off-Manifold Detergent
2020/07/20 by Christopher J. Anders, Anders, Christopher J., Pasliev, Plamen +6 · 4 citations
Computer Science · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification