Yan Teng
- Flames: Benchmarking Value Alignment of LLMs in Chinese
2023/11/12 by Kexin Huang, Huang, Kexin, Xiangyang Liu +21 · 10 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Topic Modeling
- Fake Alignment: Are LLMs Really Aligned Well?
2023/11/10 by Yixu Wang, Wang, Yixu, Yan Teng +14 · 7 citations
Computer Science · #Digital Rights Management and Security
- MEOW: MEMOry Supervised LLM Unlearning Via Inverted Facts
2024/09/18 by Tianle Gu, Kexin Huang, Gu, Tianle +10 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques
- Reflection-Bench: Evaluating Epistemic Agency in Large Language Models
2024/10/21 by Lingyu Li, Li, Lingyu, Y.D. Wang +9 · 2 citations
Computer Science · #AI-based Problem Solving and Planning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
- HoneypotNet: Backdoor Attacks Against Model Extraction
2025/01/02 by Yixu Wang, Wang, Yixu, Teng Gu +7 · 2 citations
Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning in Healthcare