Xuandong Zhao
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
2024/03/11 by Weixin Liang, Zachary Izzo, Liang, Weixin +22 · 25 voices · 43 citations
Medicine · #Artificial Intelligence in Healthcare and Education
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
2026/02/13 by Xiangyi Li, Yimin Liu, Wenbo Chen +75 · 27 voices · 24 citations
#cs.AI
- Mapping the Increasing Use of LLMs in Scientific Papers
2024/04/01 by Weixin Liang, Liang, Weixin, Yaohui Zhang +27 · 6 voices · 27 citations
Computer Science · Social Sciences · #cs.CL #cs.AI #cs.DL #cs.LG #cs.SI
- Humanity's Last Exam
2025/01/24 by Long Phan, Phan, Long, Alice Gatti +2240 · 9 voices · 103 citations
#cs.LG #cs.AI #cs.CL
- Learning to Reason without External Rewards
2025/05/26 by Xuandong Zhao, Zhao, Xuandong, Zhewei Kang +7 · 3 voices · 60 citations
#cs.LG #cs.CL
- Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models
2025/07/10 by K.S. Liang, Kaiqu Liang, Haimin Hu +10 · 15 voices · 2 citations
Social Sciences · Computer Science · #Misinformation and Its Impacts #Hate Speech and Cyberbullying Detection #Computational and Text Analysis Methods
- Provable Robust Watermarking for AI-Generated Text
2023/06/30 by Xuandong Zhao, Prabhanjan Ananth, Zhao, Xuandong +5 · 46 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Privacy-Preserving Technologies in Data #Topic Modeling
- MarkLLM: An Open-Source Toolkit for LLM Watermarking
2024/05/16 by Leyi Pan, Aiwei Liu, Pan, Leyi +20 · 31 citations
Computer Science · #Digital Rights Management and Security #Advanced Steganography and Watermarking Techniques
- Invisible Image Watermarks Are Provably Removable Using Generative AI
2023/06/02 by Xuandong Zhao, Zhao, Xuandong, Kexun Zhang +8 · 22 citations
Computer Science · #Advanced Steganography and Watermarking Techniques #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #Digital Media Forensic Detection #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis
- Scalable Best-of-N Selection for Large Language Models via Self-Certainty
2025/02/25 by Zhewei Kang, Kang, Zhewei, Xuandong Zhao +3 · 50 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- Quantifying large language model usage in scientific papers
2025/08/04 by Weixin Liang, Yaohui Zhang, Zhengxuan Wu +11 · 1 voice · 31 citations
Computer Science · Biochemistry, Genetics and Molecular Biology · #Topic Modeling #Natural Language Processing Techniques #Biomedical Text Mining and Ontologies
- Protecting Language Generation Models via Invisible Watermarking
2023/02/06 by Xuandong Zhao, Zhao, Xuandong, Yuxiang Wang +3 · 12 citations
Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Machine Learning (cs.LG)
- An Undetectable Watermark for Generative Image Models
2024/10/09 by Sam Gunn, Gunn, Sam, Xuandong Zhao +3 · 21 citations
Computer Science · #Advanced Steganography and Watermarking Techniques #Artificial Intelligence (cs.AI) #Chaos-based Image/Signal Encryption #Computer Graphics and Visualization Techniques #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM)
- Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement
2024/02/18 by Wenda Xu, Guanglei Zhu, Xu, Wenda +9 · 14 citations
Computer Science · #Topic Modeling
- DE-COP: Detecting Copyrighted Content in Language Models Training Data
2024/02/15 by André V. Duarte, Duarte, André V., Xuandong Zhao +5 · 14 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2 #Machine Learning (cs.LG) #Natural Language Processing Techniques
- Reward Shaping to Mitigate Reward Hacking in RLHF
2025/02/26 by Jiayi Fu, Xuandong Zhao, Fu, Jiayi +9 · 23 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Topic Modeling #Reinforcement Learning in Robotics
- Weak-to-Strong Jailbreaking on Large Language Models
2024/01/30 by Xuandong Zhao, Zhao, Xuandong, Xianjun Yang +11 · 10 citations
Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #Digital and Cyber Forensics #FOS: Computer and information sciences #Privacy-Preserving Technologies in Data
- A Survey on Detection of LLMs-Generated Content
2023/10/24 by Xianjun Yang, Yang, Xianjun, Liangming Pan +11 · 9 citations
Computer Science · Medicine · #Topic Modeling #Text Readability and Simplification #Artificial Intelligence in Healthcare and Education
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
2026/01/17 by Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82 · 3 voices · 4 citations
Computer Science · #cs.SE #cs.AI
- SoK: Watermarking for AI-Generated Content
2024/11/27 by Xuandong Zhao, Zhao, Xuandong, Sam Gunn +25 · 13 citations
Computer Science · #Advanced Steganography and Watermarking Techniques #Artificial Intelligence (cs.AI) #Chaos-based Image/Signal Encryption #Computer Graphics and Visualization Techniques #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
2024/02/08 by Xuandong Zhao, Zhao, Xuandong, Lei Li +3 · 6 citations
Computer Science · #Algorithms and Data Compression #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
2024/12/07 by Yuzhou Nie, Zhun Wang, Nie, Yuzhou +11 · 1 voice · 5 citations
Computer Science · #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CR #cs.LG
- Pre-trained Language Models Can be Fully Zero-Shot Learners
2022/12/14 by Xuandong Zhao, Siqi Ouyang, Zhao, Xuandong +7 · 3 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
- AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
2025/05/09 by Wang, Zhun, Vincent Siu, Zhe Ye +14 · 11 citations
Computer Science · Social Sciences · #Access Control and Trust #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Security and Verification in Computing
- Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
2025/04/07 by Tianneng Shi, Cai, Will, Xuandong Zhao +4 · 6 citations
Computer Science · Decision Sciences · #Adversarial Robustness in Machine Learning #Security and Verification in Computing #Scientific Computing and Data Management
- AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents
2025/06/17 by Xie, Jingxu, Xu, Dylan, Xuandong Zhao +3 · 9 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multi-Agent Systems and Negotiation
- Self-Sovereign Agent
2026/03/04 by Wenjie Qu, Xuandong Zhao, Jiaheng Zhang +1 · 1 voice
Computer Science · #cs.CR #cs.CY #cs.LG
- A Practical Examination of AI-Generated Text Detectors for Large Language Models
2024/12/06 by Brian Tufts, Tufts, Brian, Xuandong Zhao +3 · 2 citations
Computer Science · #Topic Modeling
- In-Context Watermarks for Large Language Models
2025/05/22 by Yepeng Liu, Xuandong Zhao, Liu, Yepeng +7 · 4 citations
Computer Science · #Adversarial Robustness in Machine Learning #Generative Adversarial Networks and Image Synthesis #Advanced Graph Neural Networks
- OVERT: A Benchmark for Over-Refusal Evaluation on Text-to-Image Models
2025/05/27 by Ziheng Cheng, Cheng, Ziheng, Yixiao Huang +11 · 2 citations
Computer Science · #Adversarial Robustness in Machine Learning