Nathan Helm-Burger
- The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
2024/03/05 by Nathaniel Li, Alexander Pan, Li, Nathaniel +114 · 2 voices · 141 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Network Security and Intrusion Detection #cs.AI #cs.CL #cs.CY #cs.LG
- Will releasing the weights of future large language models grant widespread access to pandemic agents?
2023/10/25 by Anjali Gopal, Gopal, Anjali, Nathan Helm-Burger +16 · 1 voice · 4 citations
Computer Science · Engineering · Medicine · Social Sciences · #Artificial Intelligence (cs.AI) #Biomedical and Engineering Education #FOS: Computer and information sciences #Misinformation and Its Impacts #Viral Infections and Outbreaks Research #cs.AI
- Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models
2024/12/02 by Cameron Tice, Tice, Cameron, Philipp Alexander Kreer +12 · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computational Physics and Python Applications #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Music and Audio Processing #Topic Modeling