vix.ing · top · new · best · stats · spec

Zhexin Zhang

  1. Agent-SafetyBench: Evaluating the Safety of LLM Agents
    2024/12/19 by Zhexin Zhang, Zhang, Zhexin, Shiyao Cui +11 · 2 voices · 46 citations
    Computer Science · Engineering · #Safety Systems Engineering in Autonomy #cs.CL
  2. SafetyBench: Evaluating the Safety of Large Language Models
    2023/09/13 by Zhexin Zhang, Zhang, Zhexin, Leqi Lei +16 · 30 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Text Readability and Simplification
  3. Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
    2023/11/15 by Zhexin Zhang, Zhang, Zhexin, Junxiao Yang +6 · 23 citations
    Computer Science · #Text Readability and Simplification #Natural Language Processing Techniques #Topic Modeling
  4. Unveiling the Implicit Toxicity in Large Language Models
    2023/11/29 by Jiaxin Wen, Pei Ke, Wen, Jiaxin +11 · 9 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #FOS: Computer and information sciences
  5. Ethicist: Targeted Training Data Extraction Through Loss Smoothed Soft Prompting and Calibrated Confidence Estimation
    2023/07/10 by Zhexin Zhang, Jiaxin Wen, Zhang, Zhexin +3 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  6. Knowledge-to-Jailbreak: Investigating Knowledge-driven Jailbreaking Attacks for Large Language Models
    2024/06/17 by Shangqing Tu, Zhuoran Pan, Tu, Shangqing +15 · 2 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
  7. ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
    2025/05/20 by Shiyao Cui, Qinglin Zhang, Cui, Shiyao +14 · 3 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimedia (cs.MM) #Natural Language Processing Techniques
  8. BARREL: Boundary-Aware Reasoning for Factual and Reliable LRMs
    2025/05/18 by Junxiao Yang, Yang, Junxiao, Jiyuan Tu +20 · 2 citations
    Computer Science · Decision Sciences · #Natural Language Processing Techniques #Topic Modeling #Data Quality and Management
  9. "Give a Positive Review Only": An Early Investigation Into In-Paper Prompt Injection Attacks and Defenses for AI Reviewers
    2025/11/03 by Qing Zhou, Zhou, Qin, Zhexin Zhang +5 · 2 citations
    Computer Science · #Advanced Malware Detection Techniques #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Spam and Phishing Detection
  10. Seeker: Towards Exception Safety Code Generation with Intermediate Language Agents Framework
    2024/12/16 by Xuanming Zhang, Yuxuan Chen, Zhang, Xuanming +9 · 1 citation
    Engineering · Computer Science · #Safety Systems Engineering in Autonomy #Software Reliability and Analysis Research #Software Testing and Debugging Techniques
  11. Global Challenge for Safe and Secure LLMs Track 1
    2024/11/21 by Xiaojun Jia, Yihao Huang, Jia, Xiaojun +57 · 1 citation
    Computer Science · #Law, AI, and Intellectual Property