Lin, Weiran
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
2024/03/05 by Li, Nathaniel, Pan, Alexander, Gopal, Anjali +54 · 81 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses
2024/06/07 by Weiran Lin, Anna Gerchanovsky, Lin, Weiran +9 · 4 citations
Business, Management and Accounting · #Artificial Intelligence (cs.AI) #Business Law and Ethics #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG)
- Estimating LLM Consistency: A User Baseline vs Surrogate Metrics
2025/05/26 by Wu, Xiaoyuan, Lin, Weiran, Akgul, Omer +1 · 3 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG)