vix.ing · top · new · best · stats · spec

Dombrowski, Ann-Kathrin

  1. Representation Engineering: A Top-Down Approach to AI Transparency
    2023/10/02 by Andy Zou, Zou, Andy, Long Phan +40 · 5 voices · 158 citations
    Computer Science · Engineering · #cs.LG #cs.AI #cs.CL #cs.CV #cs.CY
  2. The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
    2024/03/05 by Li, Nathaniel, Pan, Alexander, Gopal, Anjali +54 · 82 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  3. Explanations can be manipulated and geometry is to blame
    2019/06/19 by Ann-Kathrin Dombrowski, Maximilian Alber, Dombrowski, Ann-Kathrin +9 · 18 citations
    Computer Science · Decision Sciences · Medicine · #Artificial Intelligence in Healthcare and Education #Cryptography and Security (cs.CR) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Scientific Computing and Data Management
  4. Fairwashing Explanations with Off-Manifold Detergent
    2020/07/20 by Anders, Christopher J., Pasliev, Plamen, Dombrowski, Ann-Kathrin +2 · 4 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  5. Diffeomorphic Counterfactuals with Generative Models
    2022/06/10 by Dombrowski, Ann-Kathrin, Gerken, Jan E., Müller, Klaus-Robert +1 · 2 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  6. Towards Robust Explanations for Deep Neural Networks
    2020/12/18 by Dombrowski, Ann-Kathrin, Anders, Christopher J., Müller, Klaus-Robert +1 · 1 citation
    #FOS: Computer and information sciences #Machine Learning (cs.LG)