vix.ing · top · new · best · stats · spec

Wallace, Eric

  1. Scalable Extraction of Training Data from (Production) Language Models
    2023/11/28 by Milad Nasr, Nasr, Milad, Nicholas Carlini +17 · 21 voices · 66 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Topic Modeling #Explainable Artificial Intelligence (XAI)
  2. Extracting Training Data from Diffusion Models
    2023/01/30 by Nicholas Carlini, Carlini, Nicholas, Jamie Hayes +15 · 9 voices · 100 citations
    Computer Science · #Generative Adversarial Networks and Image Synthesis
  3. The False Promise of Imitating Proprietary LLMs
    2023/05/25 by Arnav Gudibande, Gudibande, Arnav, Eric Wallace +13 · 8 voices · 14 citations
    Computer Science · Engineering · #Topic Modeling #Ferroelectric and Negative Capacitance Devices #Software Engineering Research
  4. Extracting Training Data from Large Language Models
    2020/12/14 by Nicholas Carlini, Florian Tramer, Carlini, Nicholas +25 · 8 voices · 236 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Privacy-Preserving Technologies in Data #Topic Modeling #cs.CL #cs.CR #cs.LG
  5. The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
    2024/04/19 by Eric Wallace, Kai Xiao, Wallace, Eric +9 · 10 voices · 73 citations
    Computer Science · Social Sciences · #Artificial Intelligence in Law #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Legal Education and Practice Innovations #Machine Learning (cs.LG) #cs.CL #cs.CR #cs.LG
  6. Stealing Part of a Production Language Model
    2024/03/11 by Nicholas Carlini, Daniel Paleka, Carlini, Nicholas +26 · 11 voices · 26 citations
    Computer Science · #Natural Language Processing Techniques #cs.CR
  7. Large Language Models Struggle to Learn Long-Tail Knowledge
    2022/11/15 by Nikhil Kandpal, Kandpal, Nikhil, Haikang Deng +7 · 2 voices · 64 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Expert finding and Q&A systems
  8. Poisoning Language Models During Instruction Tuning
    2023/05/01 by Alexander Wan, Eric Wallace, Wan, Alexander +5 · 3 voices · 36 citations
    Computer Science · #Adversarial Robustness in Machine Learning #cs.CL #cs.CR #cs.LG
  9. GPT-4o System Card
    2024/10/25 by OpenAI, :, A. M. Hurst +503 · 1141 citations
    Medicine · #Cardiovascular Function and Risk Factors #Hyperglycemia and glycemic control in critically ill and hospitalized patients
  10. Calibrate Before Use: Improving Few-Shot Performance of Language Models
    2021/02/19 by Tony Z. Zhao, Eric Wallace, Zhao, Tony Z. +7 · 2 voices · 111 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.CL #cs.LG
  11. OpenAI o1 System Card
    2024/12/21 by OpenAI, Aaron Jaech, : +347 · 485 citations
    Computer Science · #Advanced Computational Techniques and Applications
  12. Universal Adversarial Triggers for Attacking and Analyzing NLP
    2019/08/20 by Wallace, Eric, Feng, Shi, Kandpal, Nikhil +2 · 50 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  13. InCoder: A Generative Model for Code Infilling and Synthesis
    2022/04/12 by Daniel Fried, Fried, Daniel, Armen Aghajanyan +16 · 56 citations
    Computer Science · #Software Engineering Research #Software Testing and Debugging Techniques #Software Reliability and Analysis Research
  14. gpt-oss-120b & gpt-oss-20b Model Card
    2025/08/08 by OpenAI, :, Agarwal, Sandhini +124 · 250 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
  15. Deliberative Alignment: Reasoning Enables Safer Language Models
    2024/12/20 by Guan, Melody Y., Joglekar, Manas, Wallace, Eric +12 · 71 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  16. Deduplicating Training Data Mitigates Privacy Risks in Language Models
    2022/02/14 by Kandpal, Nikhil, Wallace, Eric, Raffel, Colin · 25 citations
    #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  17. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
    2020/10/29 by Taylor Shin, Yasaman Razeghi, Shin, Taylor +7 · 16 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Sentiment Analysis and Opinion Mining
  18. Measuring Forgetting of Memorized Training Examples
    2022/06/30 by Matthew Jagielski, Om Thakkar, Jagielski, Matthew +19 · 12 citations
    Computer Science · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Privacy-Preserving Technologies in Data #Topic Modeling
  19. Do NLP Models Know Numbers? Probing Numeracy in Embeddings
    2019/09/17 by Wallace, Eric, Wang, Yizhong, Li, Sujian +2 · 8 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  20. SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore
    2023/08/08 by Sewon Min, Suchin Gururangan, Min, Sewon +11 · 1 voice · 10 citations
    Computer Science · #Topic Modeling #cs.AI #cs.CL #cs.LG
  21. What Evidence Do Language Models Find Convincing?
    2024/02/19 by Alexander Wan, Eric Wallace, Wan, Alexander +3 · 2 voices · 11 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling #cs.CL #cs.LG
  22. Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
    2024/06/28 by Halawi, Danny, Wei, Alexander, Wallace, Eric +3 · 14 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  23. Unfamiliar Finetuning Examples Control How Language Models Hallucinate
    2024/03/08 by Kang, Katie, Wallace, Eric, Tomlin, Claire +2 · 11 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  24. Cutting Down on Prompts and Parameters: Simple Few-Shot Learning with Language Models
    2021/06/24 by Robert L. Logan, Logan, Robert L., Ivana Balažević +9 · 6 citations
    Computer Science · #Topic Modeling #Domain Adaptation and Few-Shot Learning #Natural Language Processing Techniques
  25. Compositional Questions Do Not Necessitate Multi-hop Reasoning
    2019/06/07 by Sewon Min, Min, Sewon, Eric Wallace +9 · 5 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Information Retrieval and Search Behavior #Natural Language Processing Techniques #Topic Modeling
  26. Pretrained Transformers Improve Out-of-Distribution Robustness
    2020/04/13 by Hendrycks, Dan, Liu, Xiaoyuan, Wallace, Eric +3 · 5 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  27. Imitation Attacks and Defenses for Black-box Machine Translation Systems
    2020/04/30 by Eric Wallace, Wallace, Eric, Mitchell Stern +3 · 4 citations
    Computer Science · #Advanced Malware Detection Techniques #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
  28. Trading Inference-Time Compute for Adversarial Robustness
    2025/01/31 by Wojciech Zaremba, Evgenia Nitishinskaya, Zaremba, Wojciech +18 · 14 citations
    Computer Science · Engineering · #Adversarial Robustness in Machine Learning #Fault Detection and Control Systems
  29. Privacy Side Channels in Machine Learning Systems
    2023/09/11 by Edoardo Debenedetti, Debenedetti, Edoardo, Giorgio Severi +13 · 4 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Cryptography and Data Security #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Privacy-Preserving Technologies in Data
  30. AllenNLP Interpret: A Framework for Explaining Predictions of NLP Models
    2019/09/19 by Wallace, Eric, Tuyls, Jens, Wang, Junlin +3 · 2 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  31. Understanding Impacts of High-Order Loss Approximations and Features in Deep Learning Interpretation
    2019/02/01 by Sahil Singla, Singla, Sahil, Eric Wallace +4 · 4 citations
    Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Topic Modeling
  32. Estimating Worst-Case Frontier Risks of Open-Weight LLMs
    2025/08/05 by Eric Wallace, Wallace, Eric, Olivia Watkins +7 · 3 voices · 9 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.LG
  33. Detoxifying Language Models Risks Marginalizing Minority Voices
    2021/04/13 by Xu, Albert, Pathak, Eshaan, Wallace, Eric +3 · 2 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  34. Predicting Emergent Capabilities by Finetuning
    2024/11/25 by Charlie Snell, Snell, Charlie, Eric Wallace +5 · 1 voice · 3 citations
    Decision Sciences · #Complex Systems and Decision Making #cs.CL #cs.LG
  35. Trick Me If You Can: Human-in-the-loop Generation of Adversarial Examples for Question Answering
    2018/09/07 by Wallace, Eric, Rodriguez, Pedro, Feng, Shi +2 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  36. Interpreting Neural Networks With Nearest Neighbors
    2018/09/08 by Wallace, Eric, Feng, Shi, Boyd-Graber, Jordan · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  37. Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers
    2020/02/26 by Li, Zhuohan, Wallace, Eric, Shen, Sheng +4 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  38. Trustworthy AI Inference Systems: An Industry Research View
    2020/08/10 by Rosario Cammarota, Cammarota, Rosario, Matthias Schunter +41 · 1 citation
    Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #Cryptography and Data Security #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Hardware Architecture (cs.AR) #Machine Learning (cs.LG) #Privacy-Preserving Technologies in Data
  39. Evaluating Models' Local Decision Boundaries via Contrast Sets
    2020/04/06 by Matt Gardner, Gardner, Matt, Yoav Artzi +49 · 1 citation
    Computer Science · #Adversarial Robustness in Machine Learning #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Topic Modeling
  40. Automated Crossword Solving
    2022/05/19 by Wallace, Eric, Tomlin, Nicholas, Xu, Albert +4 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences