Cai, Will
- Humanity's Last Exam
2025/01/24 by Long Phan, Alice Gatti, Phan, Long +2240 · 9 voices · 106 citations
#cs.LG #cs.AI #cs.CL
- PromptArmor: Simple yet Effective Prompt Injection Defenses
2025/07/21 by Shi, Tianneng, Zhu, Kaijie, Wang, Zhun +13 · 25 citations
#Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
- Scaling Trends for Data Poisoning in LLMs
2024/08/06 by Dillon Bowen, Brendan Murphy, Bowen, Dillon +9 · 7 citations
Computer Science · Decision Sciences · #Privacy-Preserving Technologies in Data #Scientific Computing and Data Management
- Improving LLM Safety Alignment with Dual-Objective Optimization
2025/03/05 by Zhao, Xuandong, Cai, Will, Tingyan Shi +7 · 10 citations
Computer Science · #Adversarial Robustness in Machine Learning #Topic Modeling #Explainable Artificial Intelligence (XAI)
- Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
2025/04/07 by Cai, Will, Tianneng Shi, Xuandong Zhao +4 · 8 citations
Computer Science · Decision Sciences · #Adversarial Robustness in Machine Learning #Security and Verification in Computing #Scientific Computing and Data Management
- The Geometry of Harmfulness in LLMs through Subconcept Probing
2025/07/23 by Shah, McNair, Angeline, Saleena, Kumar, Adhitya Rajendra +5 · 3 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences