Jia, Ruoxi
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
2023/10/05 by Xiangyu Qi, Qi, Xiangyu, Yi Zeng +11 · 165 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
2024/01/12 by Yi Zeng, Hongpeng Lin, Zeng, Yi +9 · 94 citations
Social Sciences · Computer Science · #Ethics and Social Impacts of AI #Hate Speech and Cyberbullying Detection #Misinformation and Its Impacts
- Towards Efficient Data Valuation Based on the Shapley Value
2019/02/27 by Ruoxi Jia, Jia, Ruoxi, David Dao +17 · 37 citations
Computer Science · Decision Sciences · Economics, Econometrics and Finance · #Auction Theory and Applications #Blockchain Technology Applications and Security #FOS: Computer and information sciences #Game Theory and Voting Systems #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models
2023/08/20 by Bilgehan Sel, Ahmad Al-Tawaha, Sel, Bilgehan +7 · 1 voice · 8 citations
Computer Science · Decision Sciences · #cs.CL #cs.AI
- SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
2024/06/20 by Tinghao Xie, Xie, Tinghao, Xiangyu Qi +29 · 39 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Natural Language Processing Techniques #Software Reliability and Analysis Research #Topic Modeling
- A Principled Approach to Data Valuation for Federated Learning
2020/09/14 by Tianhao Wang, Johannes Rausch, Wang, Tianhao +7 · 17 citations
Computer Science · #Privacy-Preserving Technologies in Data #Adversarial Robustness in Machine Learning #Domain Adaptation and Few-Shot Learning
- Narcissus: A Practical Clean-Label Backdoor Attack with Limited Information
2022/04/11 by Yi Zeng, Zeng, Yi, Minzhou Pan +9 · 17 citations
Computer Science · #Advanced Malware Detection Techniques #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- The Secret Revealer: Generative Model-Inversion Attacks Against Deep Neural Networks
2019/11/17 by Yuheng Zhang, Zhang, Yuheng, Ruoxi Jia +9 · 10 citations
Computer Science · Engineering · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Geophysical Methods and Applications #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Privacy-Preserving Technologies in Data
- AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies
2024/07/11 by Yi Zeng, Zeng, Yi, Yu Yang +21 · 1 voice · 18 citations
#cs.CY #cs.AI
- Data Shapley in One Training Run
2024/06/16 by Jiachen T. Wang, Wang, Jiachen T., Prateek Mittal +5 · 20 citations
Computer Science · #Computation and Language (cs.CL) #Edcuational Technology Systems #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Teaching and Learning Programming
- A Safe Harbor for AI Evaluation and Red Teaming
2024/03/07 by Shayne Longpre, Sayash Kapoor, Longpre, Shayne +43 · 1 voice · 15 citations
Computer Science · #Explainable Artificial Intelligence (XAI)
- Adversarial Unlearning of Backdoors via Implicit Hypergradient
2021/10/07 by Yi Zeng, Si Chen, Zeng, Yi +9 · 10 citations
Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Advanced Malware Detection Techniques
- Knowledge-Enriched Distributional Model Inversion Attacks
2020/10/08 by Chen, Si, Kahla, Mostafa, Jia, Ruoxi +1 · 7 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content
2024/03/19 by Yuan, Zhuowen, Xiong, Zidi, Zeng, Yi +4 · 13 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Rethinking the Backdoor Attacks' Triggers: A Frequency Perspective
2021/04/07 by Zeng, Yi, Park, Won, Mao, Z. Morley +1 · 7 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models
2024/06/24 by Yi Zeng, Zeng, Yi, Weiyu Sun +9 · 13 citations
Computer Science · #Adversarial Robustness in Machine Learning #Software Testing and Debugging Techniques
- LAVA: Data Valuation without Pre-Specified Learning Algorithms
2023/04/28 by Hoang Anh Just, Feiyang Kang, Just, Hoang Anh +11 · 9 citations
Computer Science · #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Privacy-Preserving Technologies in Data #Stochastic Gradient Optimization Techniques
- Selective Differential Privacy for Language Modeling
2021/08/30 by Weiyan Shi, Aiqi Cui, Shi, Weiyan +7 · 6 citations
Computer Science · Social Sciences · #Access Control and Trust #Computation and Language (cs.CL) #Cryptography and Data Security #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Privacy-Preserving Technologies in Data
- Label-Only Model Inversion Attacks via Boundary Repulsion
2022/03/03 by Kahla, Mostafa, Chen, Si, Just, Hoang Anh +1 · 6 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Improving Robustness to Model Inversion Attacks via Mutual Information Regularization
2020/09/11 by Tianhao Wang, Wang, Tianhao, Yuheng Zhang +3 · 5 citations
Computer Science · #Adversarial Robustness in Machine Learning #Cryptography and Data Security #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Privacy-Preserving Technologies in Data
- InfoBERT: Improving Robustness of Language Models from An Information\n Theoretic Perspective
2020/10/05 by Boxin Wang, Shuohang Wang, Wang, Boxin +11 · 5 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- Just Fine-tune Twice: Selective Differential Privacy for Large Language Models
2022/04/15 by Shi, Weiyan, Shea, Ryan, Chen, Si +3 · 5 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences
- Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety
2025/02/02 by Xingjun Ma, Ma, Xingjun, Yifeng Gao +88 · 20 citations
Health Professions · Decision Sciences · Engineering · #Occupational Health and Safety Research #Risk and Safety Analysis #Safety Systems Engineering in Autonomy
- Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMs
2024/05/05 by Feiyang Kang, Hoang Anh Just, Kang, Feiyang +13 · 8 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification #Mineral Processing and Grinding #Natural Language Processing Techniques
- Private Data Valuation and Fair Payment in Data Marketplaces
2022/10/17 by Tian, Zhihua, Liu, Jian, Li, Jingyu +5 · 5 citations
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences
- AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies
2024/06/25 by Zeng, Yi, Klyman, Kevin, Zhou, Andy +6 · 8 citations
#Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #FOS: Computer and information sciences
- Variance reduced Shapley value estimation for trustworthy data valuation
2022/10/30 by Mengmeng Wu, Ruoxi Jia, Wu, Mengmeng +7 · 5 citations
Computer Science · Decision Sciences · #Data Quality and Management #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Privacy-Preserving Technologies in Data #Probability and Risk Models
- AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
2024/07/29 by Feiyang Kang, Yifan Sun, Kang, Feiyang +11 · 7 citations
Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Data Mining Algorithms and Applications #Data Quality and Management #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Semantic Web and Ontologies
- Threshold KNN-Shapley: A Linear-Time and Privacy-Friendly Approach to Data Valuation
2023/08/30 by Wang, Jiachen T., Zhu, Yuqing, Wang, Yu-Xiang +2 · 5 citations
#Computer Science and Game Theory (cs.GT) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Practical Membership Inference Attacks Against Large-Scale Multi-Modal Models: A Pilot Study
2023/09/29 by Myeongseob Ko, Ko, Myeongseob, Ming Jin +5 · 6 citations
Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications
- JIGMARK: A Black-Box Approach for Enhancing Image Watermarks against Diffusion Model Edits
2024/06/06 by Minzhou Pan, Pan, Minzhou, Yi Zeng +11 · 5 citations
Computer Science · #Advanced Steganography and Watermarking Techniques #Computer Graphics and Visualization Techniques #Digital Media Forensic Detection
- Robust Anomaly Detection and Backdoor Attack Detection Via Differential Privacy
2019/11/16 by Min Du, Ruoxi Jia, Du, Min +3 · 5 citations
Computer Science · #Anomaly Detection Techniques and Applications #Adversarial Robustness in Machine Learning #Privacy-Preserving Technologies in Data
- CATER: Intellectual Property Protection on Text Generation APIs via Conditional Watermarks
2022/09/19 by Xuanli He, Qiongkai Xu, He, Xuanli +11 · 3 citations
Computer Science · #Advanced Malware Detection Techniques #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Internet Traffic Analysis and Secure E-voting #Security and Verification in Computing
- Capturing the Temporal Dependence of Training Data Influence
2024/12/12 by Wang, Jiachen T., Song, Dawn, Zou, James +2 · 7 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- LLM Can be a Dangerous Persuader: Empirical Study of Persuasion Safety in Large Language Models
2025/04/14 by Minqian Liu, Liu, Minqian, Zhiyang Xu +19 · 8 citations
Computer Science · Social Sciences · #Topic Modeling #Ethics and Social Impacts of AI #Hate Speech and Cyberbullying Detection
- DPlis: Boosting Utility of Differentially Private Deep Learning via Randomized Smoothing
2021/03/02 by Wenxiao Wang, Wang, Wenxiao, Tianhao Wang +11 · 3 citations
Computer Science · #Privacy-Preserving Technologies in Data #Adversarial Robustness in Machine Learning #Stochastic Gradient Optimization Techniques
- Rethinking Data Shapley for Data Selection Tasks: Misleads and Merits
2024/05/06 by Jiachen T. Wang, T.S. Yang, Wang, Jiachen T. +7 · 4 citations
Computer Science · #Data Mining Algorithms and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- On Solution Functions of Optimization: Universal Approximation and Covering Number Bounds
2022/12/02 by Jin, Ming, Khattar, Vanshaj, Kaushik, Harshal +2 · 2 citations
#FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Optimization and Control (math.OC)
- Boosting Alignment for Post-Unlearning Text-to-Image Generative Models
2024/12/09 by Ko, Myeongseob, Li, Henry, Wang, Zhun +6 · 5 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- ASSET: Robust Backdoor Data Detection Across a Multiplicity of Deep Learning Paradigms
2023/02/22 by Pan, Minzhou, Zeng, Yi, Lyu, Lingjuan +2 · 2 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- 2D-Shapley: A Framework for Fragmented Data Valuation
2023/06/18 by Zhihong Liu, Hoang Anh Just, Liu, Zhihong +7 · 2 citations
Computer Science · Decision Sciences · #Bayesian Modeling and Causal Inference #Data Quality and Management #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Learning to Rank for Active Learning via Multi-Task Bilevel Optimization
2023/10/25 by Zixin Ding, Ding, Zixin, Si Chen +5 · 2 citations
Computer Science · Engineering · #Machine Learning and Algorithms #Machine Learning and Data Classification #Reservoir Engineering and Simulation Methods
- Strategize Globally, Adapt Locally: A Multi-Turn Red Teaming Agent with Dual-Level Learning
2025/04/02 by Chen, Si, Yu, Xiao, Mehrabi, Ninareh +3 · 5 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
- Efficient Data Shapley for Weighted Nearest Neighbor Algorithms
2024/01/20 by Wang, Jiachen T., Mittal, Prateek, Jia, Ruoxi · 2 citations
#Data Structures and Algorithms (cs.DS) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Scalability vs. Utility: Do We Have to Sacrifice One for the Other in Data Importance Quantification?
2019/11/17 by Ruoxi Jia, Fan Wu, Jia, Ruoxi +15 · 1 citation
Computer Science · #Data Stream Mining Techniques #FOS: Computer and information sciences #Imbalanced Data Classification Techniques #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification
- A Survey on Data Markets
2024/11/09 by Zhang, Jiayao, Bi, Yuran, Cheng, Mengye +15 · 3 citations
#Artificial Intelligence (cs.AI) #Computer Science and Game Theory (cs.GT) #Databases (cs.DB) #FOS: Computer and information sciences
- Data-Centric Human Preference with Rationales for Direct Preference Alignment
2024/07/19 by Hoang Anh Just, Just, Hoang Anh, Ming Jin +7 · 2 citations
Decision Sciences · Computer Science · #Multi-Criteria Decision Making #Data Management and Algorithms
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
2025/10/02 by Feiyang Kang, Newsha Ardalani, Kang, Feiyang +17 · 3 citations
Social Sciences · #Artificial Intelligence in Law
- Improving Cooperative Game Theory-based Data Valuation via Data Utility Learning
2021/07/13 by Wang, Tianhao, Yang, Yu, Jia, Ruoxi · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models
2024/11/15 by Jianhong Tu, Zheng Ni, Tu, Jianhong +19 · 2 citations
Computer Science · #Speech and dialogue systems #Natural Language Processing Techniques
- Performance Scaling via Optimal Transport: Enabling Data Selection from Partially Revealed Sources
2023/07/05 by Kang, Feiyang, Just, Hoang Anh, Sahu, Anit Kumar +1 · 2 citations
#Artificial Intelligence (cs.AI) #Computational Engineering #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Finance #Machine Learning (cs.LG) #and Science (cs.CE)
- LLMs Can Plan Only If We Tell Them
2025/01/23 by Bilgehan Sel, Sel, Bilgehan, Jin Ming +2 · 3 citations
Biochemistry, Genetics and Molecular Biology · #Cancer Genomics and Diagnostics
- Revisiting Data-Free Knowledge Distillation with Poisoned Teachers
2023/06/04 by Hong, Junyuan, Zeng, Yi, Yu, Shuyang +3 · 1 citation
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Who Leaked the Model? Tracking IP Infringers in Accountable Federated Learning
2023/12/06 by Yu, Shuyang, Hong, Junyuan, Zeng, Yi +3 · 1 citation
#Cryptography and Security (cs.CR) #FOS: Computer and information sciences
- Data Acquisition: A New Frontier in Data-centric AI
2023/11/22 by Lingjiao Chen, Bilge Acun, Chen, Lingjiao +19 · 1 citation
Decision Sciences · Business, Management and Accounting · Computer Science · #Data Quality and Management #Big Data and Business Intelligence #Privacy-Preserving Technologies in Data
- Privacy-Enhanced Architecture for Occupancy-based HVAC Control
2016/07/11 by Ruoxi Jia, Jia, Ruoxi, Roy Dong +5 · 1 citation
Engineering · Medicine · Psychology · #Building Energy and Comfort Optimization #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #FOS: Electrical engineering #Facilities and Workplace Management #Healthcare Technology and Patient Monitoring #Systems and Control (eess.SY) #electronic engineering #information engineering
- Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
2025/07/06 by Dabas, Mahavir, Chen, Si, Fleming, Charles +2 · 2 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Detecting Adversarial Data via Perturbation Forgery
2024/05/25 by Qian Wang, Chen Li, Wang, Qian +12 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Computer Vision and Pattern Recognition (cs.CV) #Digital Media Forensic Detection #FOS: Computer and information sciences #Machine Learning (cs.LG)
- AI Risk Management Should Incorporate Both Safety and Security
2024/05/29 by Xiangyu Qi, Yangsibo Huang, Qi, Xiangyu +47 · 1 citation
Medicine · Social Sciences · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Healthcare and Education #Cryptography and Security (cs.CR) #Ethics and Social Impacts of AI #FOS: Computer and information sciences
- Can We Trust the Performance Evaluation of Uncertainty Estimation Methods in Text Summarization?
2024/06/25 by He, Jianfeng, Yang, Runing, Yu, Linlin +5 · 1 citation
#Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Skin-in-the-Game: Decision Making via Multi-Stakeholder Alignment in LLMs
2024/05/21 by Sel, Bilgehan, Shanmugasundaram, Priya, Kachuee, Mohammad +3 · 1 citation
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- MAViS: A Multi-Agent Framework for Long-Sequence Video Storytelling
2025/08/11 by Qian Wang, Wang, Qian, Huang, Ziqi +5 · 3 citations
Computer Science · Engineering · #Artificial Intelligence in Games #Video Analysis and Summarization #Human Motion and Animation
- AdaDeDup: Adaptive Hybrid Data Pruning for Efficient Large-Scale Object Detection Training
2025/06/24 by Kang, Feiyang, Chang, Nadine, Shen, Maying +4 · 1 citation
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- DiPT: Enhancing LLM reasoning through diversified perspective-taking
2024/09/10 by Just, Hoang Anh, Mahavir Dabas, Dabas, Mahavir +6 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques
- FASTTRACK: Fast and Accurate Fact Tracing for LLMs
2024/04/22 by Si Chen, Chen, Si, Feiyang Kang +5 · 1 citation
Decision Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Data Quality and Management #FOS: Computer and information sciences
- Probing Knowledge Holes in Unlearned LLMs
2025/10/27 by Myeongseob Ko, Hoang Anh Just, Ko, Myeongseob +6 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Domain Adaptation and Few-Shot Learning #Generative Adversarial Networks and Image Synthesis