Mansi Phute
- LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
2023/08/14 by Mansi Phute, Alec Helbling, Phute, Mansi +8 · 38 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Adversarial Robustness in Machine Learning
- Semi-Truths: A Large-Scale Dataset of AI-Augmented Images for Evaluating Robustness of AI-Generated Image detectors
2024/11/12 by Anisha Pal, Julia Kruk, Pal, Anisha +11 · 1 voice · 4 citations
Computer Science · Medicine · #Adversarial Robustness in Machine Learning #COVID-19 diagnosis using AI
- Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
2025/06/05 by Seongmin Lee, Aeree Cho, Lee, Seongmin +9 · 3 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Software Engineering (cs.SE) #Topic Modeling
- ALLUDE: A Unified Evaluation System for Configurable Attacks in Differentiable Environments
2026/07/19 by Mansi Phute, Alexander Greenhalgh, Matthew Hull +8
#cs.CV #cs.AI