vix.ing · top · new · best · stats · spec

Julian Minder

  1. Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
    2025/04/03 by Julian Minder, Minder, Julian, Clément Dumas +7 · 3 voices · 5 citations
    #cs.LG #cs.AI #cs.CL
  2. Controllable Context Sensitivity and the Knob Behind It
    2024/11/11 by Julian Minder, Kevin Du, Minder, Julian +11 · 1 voice · 10 citations
    Computer Science · #Explainable Artificial Intelligence (XAI) #Multimodal Machine Learning Applications #Topic Modeling #cs.AI #cs.CL
  3. The Non-Linear Representation Dilemma: Is Causal Abstraction Enough for Mechanistic Interpretability?
    2025/07/11 by Denis Sutter, Julian Minder, Sutter, Denis +5 · 3 voices · 9 citations
    Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.LG
  4. Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
    2025/10/14 by Julian Minder, Minder, Julian, Clément Dumas +11 · 1 voice · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.AI #cs.CL
  5. Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
    2025/12/17 by Adam Karvonen, Karvonen, Adam, James Chua +19 · 1 voice · 1 citation
    Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Explainable Artificial Intelligence (XAI) #Topic Modeling #cs.AI #cs.CL #cs.LG