Lin, Leon
- Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
2024/07/03 by Brown, Hannah, Lin, Leon, Kawaguchi, Kenji +1 · 4 citations
#Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Single Character Perturbations Break LLM Alignment
2024/07/03 by Leon Lin, H. Alex Brown, Lin, Leon +5 · 3 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques