vix.ing · top · new · best · stats · spec

Ziser, Yftah

  1. Spectral Editing of Activations for Large Language Model Alignment
    2024/05/15 by Yifu Qiu, Zheng Zhao, Qiu, Yifu +9 · 9 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
  2. Detecting and Mitigating Hallucinations in Multilingual Summarisation
    2023/05/23 by Qiu, Yifu, Ziser, Yftah, Korhonen, Anna +2 · 5 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  3. Are Large Language Models Temporally Grounded?
    2023/11/14 by Qiu, Yifu, Zhao, Zheng, Ziser, Yftah +3 · 5 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
  4. Layer by Layer: Uncovering Where Multi-Task Learning Happens in Instruction-Tuned Large Language Models
    2024/10/25 by Zhao, Zheng, Ziser, Yftah, Cohen, Shay B. · 3 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  5. SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs
    2025/06/01 by Ghosh, Shaona, Bhattacharjee, Amrita, Ziser, Yftah +1 · 6 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG)
  6. Knowing Before Saying: LLM Representations Encode Information About Chain-of-Thought Success Before Completion
    2025/05/30 by Afzal, Anum, Matthes, Florian, Chechik, Gal +1 · 4 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  7. Beyond Next Token Probabilities: Learnable, Fast Detection of Hallucinations and Data Contamination on LLM Output Distributions
    2025/03/18 by Bar-Shalom, Guy, Frasca, Fabrizio, Lim, Derek +5 · 4 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG)
  8. Gold Doesn't Always Glitter: Spectral Removal of Linear and Nonlinear Guarded Attribute Information
    2022/03/15 by Shao, Shun, Ziser, Yftah, Cohen, Shay B. · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  9. Erasure of Unaligned Attributes from Neural Representations
    2023/02/06 by Shao, Shun, Ziser, Yftah, Cohen, Shay · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences