vix.ing · top · new · best · stats · spec

Miles Turpin

  1. Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
    2023/05/07 by Miles Turpin, Julian Michael, Turpin, Miles +5 · 209 citations
    Computer Science · Materials Science · #Topic Modeling #Explainable Artificial Intelligence (XAI) #Machine Learning in Materials Science
  2. Looking Inward: Language Models Can Learn About Themselves by Introspection
    2024/10/17 by Felix J Binder, Binder, Felix J, James Chua +15 · 4 voices · 36 citations
    Computer Science · #Natural Language Processing Techniques
  3. Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought
    2024/03/08 by James Chua, Chua, James, Edward Rees +11 · 6 citations
    Decision Sciences · Computer Science · #Decision-Making and Behavioral Economics #AI in Service Interactions
  4. Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
    2025/06/28 by Miles Turpin, Turpin, Miles, Man Li +6 · 7 citations
    Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences