vix.ing · top · new · best · stats · spec

Niloofar Mireshghallah

  1. Alignment Whack-a-Mole : Finetuning Activates Verbatim Recall of Copyrighted Books in Large Language Models
    2026/03/21 by Xinyue Liu, Niloofar Mireshghallah, Jane C. Ginsburg +1 · 23 voices · 2 citations
    #cs.CL #cs.AI #cs.CY
  2. Learning to Reason in 13 Parameters
    2026/02/04 by John X. Morris, Niloofar Mireshghallah, Mark Ibrahim +1 · 13 voices · 2 citations
    #cs.LG
  3. AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
    2024/10/05 by Ximing Lu, Lu, Ximing, Melanie Sclar +19 · 8 voices · 13 citations
    #cs.CL
  4. A Roadmap to Pluralistic Alignment
    2024/02/07 by Taylor Sorensen, Jared Moore, Sorensen, Taylor +21 · 1 voice · 30 citations
    Social Sciences · Computer Science · Medicine · #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #Artificial Intelligence in Healthcare and Education
  5. Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
    2023/10/27 by Niloofar Mireshghallah, Mireshghallah, Niloofar, Hyunwoo Kim +11 · 41 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Privacy-Preserving Technologies in Data #Topic Modeling
  6. WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
    2024/06/26 by Liwei Jiang, Jiang, Liwei, Kavel Rao +19 · 49 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #FOS: Computer and information sciences
  7. Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
    2024/12/09 by A. Feder Cooper, Christopher A. Choquette-Choo, Cooper, A. Feder +75 · 5 voices · 8 citations
    Computer Science · Social Sciences · #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.CY #cs.LG
  8. Do Membership Inference Attacks Work on Large Language Models?
    2024/02/12 by Michael Duan, Anshuman Suri, Duan, Michael +17 · 29 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Topic Modeling
  9. Trust No Bot: Discovering Personal Disclosures in Human-LLM Conversations in the Wild
    2024/07/16 by Niloofar Mireshghallah, Mireshghallah, Niloofar, Maria Antoniak +7 · 2 voices · 21 citations
    Computer Science · Social Sciences · #Cybercrime and Law Enforcement Studies #Privacy, Security, and Data Protection #Spam and Phishing Detection #cs.CL
  10. Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language Models
    2025/03/15 by Abhilasha Ravichander, Ravichander, Abhilasha, Jillian Fisher +15 · 1 voice · 5 citations
    #cs.CL
  11. Privacy Ripple Effects from Adding or Removing Personal Information in Language Model Training
    2025/02/21 by Jaydeep Borkar, Matthew Jagielski, Borkar, Jaydeep +9 · 3 voices · 5 citations
    #cs.CL #cs.CR
  12. CopyBench: Measuring Literal and Non-Literal Reproduction of Copyright-Protected Text in Language Model Generation
    2024/07/09 by Tong Chen, Chen, Tong, Akari Asai +15 · 6 citations
    Computer Science · #Computation and Language (cs.CL) #Digital Rights Management and Security #FOS: Computer and information sciences #Machine Learning (cs.LG)
  13. Alpaca against Vicuna: Using LLMs to Uncover Memorization of LLMs
    2024/03/05 by Aly M. Kassem, Omar Mahmoud, Kassem, Aly M. +13 · 4 citations
    Social Sciences · #Artificial Intelligence in Law
  14. Misusing Tools in Large Language Models With Visual Adversarial Examples
    2023/10/04 by Xiaohan Fu, Zihan Wang, Fu, Xiaohan +11 · 3 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Topic Modeling
  15. Exploring the limits of strong membership inference attacks on large language models
    2025/05/24 by Jamie Hayes, Ilia Shumailov, Hayes, Jamie +29 · 2 voices · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling #cs.AI #cs.CR #cs.LG
  16. Spectrum Tuning: Post-Training for Distributional Coverage and In-Context Steerability
    2025/10/07 by Taylor Sorensen, Sorensen, Taylor, Benjamin T. Newman +13 · 7 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Blind Source Separation Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences
  17. Leaky Language Models: Stealing Architecture and Inference Optimizations via Per-Token Timing
    2026/07/22 by Sadegh Majidi, Niloofar Mireshghallah, Kazem Taram · 1 voice
    #cs.CR #cs.LG
  18. Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
    2025/01/20 by Ali Naseh, Naseh, Ali, Niloofar Mireshghallah +1 · 2 citations
    Computer Science · Social Sciences · #Natural Language Processing Techniques #Topic Modeling #Computational and Text Analysis Methods
  19. A False Sense of Privacy: Evaluating Textual Data Sanitization Beyond Surface-level Privacy Leakage
    2025/04/28 by Rui Xin, Niloofar Mireshghallah, Xin, Rui +15 · 2 citations
    Computer Science · Social Sciences · #Advanced Malware Detection Techniques #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Privacy, Security, and Data Protection #Privacy-Preserving Technologies in Data