vix.ing · top · new · best · stats · spec

Zhang, Wenhui

  1. Introducing v0.5 of the AI Safety Benchmark from MLCommons
    2024/04/18 by Bertie Vidgen, Vidgen, Bertie, Adarsh Agrawal +202 · 2 voices · 10 citations
    Computer Science · #Adversarial Robustness in Machine Learning
  2. RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
    2024/07/23 by Xu, Huiyu, Zhang, Wenhui, Wang, Zhibo +5 · 10 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
  3. JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit
    2024/11/17 by He, Zeqing, Wang, Zhibo, Chu, Zhixuan +4 · 8 citations
    #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
  4. CXXCrafter: An LLM-Based Agent for Automated C/C++ Open Source Software Building
    2025/05/27 by Yu, Zhengmin, Zhang, Yuan, Wen, Ming +3 · 6 citations
    #FOS: Computer and information sciences #Software Engineering (cs.SE)
  5. Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
    2025/08/11 by Chu, Kexin, Lin, Zecheng, Dawei Xiang +15 · 7 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Operating Systems (cs.OS) #Parallel Computing and Optimization Techniques #Security and Verification in Computing
  6. ZTaint-Havoc: From Havoc Mode to Zero-Execution Fuzzing-Driven Taint Inference
    2025/06/10 by Yuchong Xie, Wenhui Zhang, Xie, Yuchong +3 · 1 citation
    Computer Science · #Advanced Malware Detection Techniques #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Security and Verification in Computing #Software Engineering (cs.SE) #Software Testing and Debugging Techniques
  7. Interpretable LLM Guardrails via Sparse Representation Steering
    2025/03/21 by He, Zeqing, Wang, Zhibo, Xu, Huiyu +3 · 1 citation
    #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
  8. Can Small Language Models Reliably Resist Jailbreak Attacks? A Comprehensive Evaluation
    2025/03/09 by Zhang, Wenhui, Xu, Huiyu, Wang, Zhibo +3 · 1 citation
    #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences