Xu, Huiyu
- RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
2024/07/23 by Xu, Huiyu, Zhang, Wenhui, Wang, Zhibo +5 · 10 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
- JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit
2024/11/17 by He, Zeqing, Wang, Zhibo, Chu, Zhixuan +4 · 8 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences
- Interpretable LLM Guardrails via Sparse Representation Steering
2025/03/21 by He, Zeqing, Wang, Zhibo, Xu, Huiyu +3 · 1 citation
#Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
- Can Small Language Models Reliably Resist Jailbreak Attacks? A Comprehensive Evaluation
2025/03/09 by Zhang, Wenhui, Xu, Huiyu, Wang, Zhibo +3 · 1 citation
#Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences