Himabindu Lakkaraju
- Manipulating Large Language Models to Increase Product Visibility
2024/04/11 by Aounon Kumar, Kumar, Aounon, Himabindu Lakkaraju +1 · 8 voices · 8 citations
Computer Science · #Topic Modeling #Recommender Systems and Techniques #Expert finding and Q&A systems
- Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE)
2024/02/16 by Usha Bhalla, Alex Oesterling, Bhalla, Usha +8 · 2 voices · 28 citations
Computer Science · #Bayesian Modeling and Causal Inference #cs.CV #cs.LG
- Fooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation\n Methods
2019/11/06 by Dylan Slack, Sophie Hilgard, Slack, Dylan +7 · 37 citations
Computer Science · Medicine · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Artificial Intelligence (cs.AI) #Artificial Intelligence in Healthcare and Education #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
2024/02/07 by Chirag Agarwal, Agarwal, Chirag, Sree Harsha Tanneru +3 · 2 voices · 21 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.CL
- Reliable Post hoc Explanations: Modeling Uncertainty in Explainability
2020/08/11 by Dylan Slack, Sophie Hilgard, Slack, Dylan +5 · 11 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Machine Learning and Data Classification #Adversarial Robustness in Machine Learning
- The Disagreement Problem in Explainable Machine Learning: A Practitioner's Perspective
2022/02/03 by Satyapriya Krishna, Krishna, Satyapriya, Tessa Han +9 · 9 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #Machine Learning and Data Classification #Adversarial Robustness in Machine Learning
- Counterfactual Explanations Can Be Manipulated
2021/06/04 by Dylan Slack, Sophie Hilgard, Slack, Dylan +5 · 7 citations
Computer Science · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Imbalanced Data Classification Techniques #Machine Learning (cs.LG)
- "How do I fool you?": Manipulating User Trust via Misleading Black Box Explanations
2019/11/15 by Himabindu Lakkaraju, Lakkaraju, Himabindu, Osbert Bastani +1 · 6 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Privacy-Preserving Technologies in Data
- Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
2024/02/27 by Zhenting Qi, Qi, Zhenting, Hanlin Zhang +7 · 12 citations
Computer Science · #Algorithms and Data Compression #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Semantic Web and Ontologies #Web Data Mining and Analysis
- Towards Robust and Reliable Algorithmic Recourse
2021/02/26 by Sohini Upadhyay, Upadhyay, Sohini, Shalmali Joshi +3 · 6 citations
Computer Science · Medicine · #Explainable Artificial Intelligence (XAI) #Adversarial Robustness in Machine Learning #Artificial Intelligence in Healthcare and Education
- TalkToModel: Explaining Machine Learning Models with Interactive Natural Language Conversations
2022/07/08 by Dylan Slack, Slack, Dylan, Satyapriya Krishna +5 · 7 citations
Computer Science · Medicine · #Explainable Artificial Intelligence (XAI) #Artificial Intelligence in Healthcare and Education #Topic Modeling
- Which Explanation Should I Choose? A Function Approximation Perspective to Characterizing Post Hoc Explanations
2022/06/02 by Tessa Han, Han, Tessa, Suraj Srinivas +3 · 6 citations
Computer Science · Decision Sciences · #Explainable Artificial Intelligence (XAI) #Data Stream Mining Techniques #Advanced Bandit Algorithms Research
- Towards the Unification and Robustness of Perturbation and Gradient Based Explanations
2021/02/21 by Sushant Agarwal, Agarwal, Sushant, Shahin Jabbari +10 · 1 voice · 4 citations
Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Healthcare #cs.LG
- Towards Uncovering How Large Language Model Works: An Explainability Perspective
2024/02/16 by Haiyan Zhao, Zhao, Haiyan, Fan Yang +6 · 8 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
- Rethinking Stability for Attribution-based Explanations
2022/03/14 by Chirag Agarwal, Agarwal, Chirag, Nari Johnson +10 · 5 citations
Computer Science · Decision Sciences · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Scientific Computing and Data Management
- Fair Influence Maximization: A Welfare Optimization Approach
2020/06/14 by Aida Rahmattalabi, Rahmattalabi, Aida, Shahin Jabbari +13 · 4 citations
Mathematics · #Advanced Causal Inference Techniques
- OpenXAI: Towards a Transparent Evaluation of Model Explanations
2022/06/22 by Chirag Agarwal, Agarwal, Chirag, Ley, Dan +14 · 5 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Healthcare
- Analyzing Chain-of-Thought Prompting in Large Language Models via Gradient-based Feature Attributions
2023/07/25 by Skyler Wu, Eric Meng Shen, Wu, Skyler +7 · 6 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #Expert finding and Q&A systems #FOS: Computer and information sciences #Topic Modeling
- Probing GNN Explainers: A Rigorous Theoretical and Empirical Analysis of GNN Explanation Methods
2021/06/16 by Chirag Agarwal, Marinka Žitnik, Agarwal, Chirag +3 · 4 citations
Computer Science · #Advanced Graph Neural Networks #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Healthcare
- On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language Models
2024/06/15 by Sree Harsha Tanneru, Tanneru, Sree Harsha, Dan Ley +5 · 8 citations
Computer Science · #Advanced Graph Neural Networks #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
- Interpretability Needs a New Paradigm
2024/05/08 by Andreas Madsen, Andreas Nygaard Madsen, Madsen, Andreas +6 · 1 voice · 4 citations
Health Professions · #Interpreting and Communication in Healthcare
- Detecting LLM-Generated Peer Reviews
2025/03/20 by Vishisht Rao, Aounon Kumar, Rao, Vishisht +5 · 4 voices · 4 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #Digital Libraries (cs.DL) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Spam and Phishing Detection #cs.AI #cs.CR #cs.DL
- Does Fair Ranking Improve Minority Outcomes? Understanding the Interplay\n of Human and Algorithmic Biases in Online Hiring
2020/12/01 by Tom Sühr, Sühr, Tom, Sophie Hilgard +3 · 2 citations
Social Sciences · Economics, Econometrics and Finance · Business, Management and Accounting · #Names, Identity, and Discrimination Research #Game Theory and Voting Systems #Employer Branding and e-HRM
- On the Impact of Fine-Tuning on Chain-of-Thought Reasoning
2024/11/22 by Elita Lobo, Lobo, Elita, Chirag Agarwal +3 · 5 citations
Computer Science · #Advanced Text Analysis Techniques #Cognitive Science and Mapping #Computation and Language (cs.CL) #FOS: Computer and information sciences
- Rethinking Explainability as a Dialogue: A Practitioner's Perspective
2022/02/03 by Himabindu Lakkaraju, Dylan Slack, Lakkaraju, Himabindu +7 · 2 citations
Business, Management and Accounting · Computer Science · #Business Process Modeling and Analysis #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Topic Modeling
- Word-Level Explanations for Analyzing Bias in Text-to-Image Models
2023/06/03 by Alexander Y. Lin, Lucas Monteiro Paes, Lin, Alexander +7 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models
2025/05/19 by Zidi Xiong, Xiong, Zidi, Shan Chen +5 · 6 citations
Computer Science · #Artificial Intelligence (cs.AI) #Constraint Satisfaction and Optimization #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Topic Modeling
- In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
2023/10/09 by Nicholas Kroeger, Kroeger, Nicholas, Dan Ley +7 · 2 citations
Computer Science · Materials Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning in Materials Science #Topic Modeling
- Quantifying Uncertainty in Natural Language Explanations of Large Language Models
2023/11/06 by Sree Harsha Tanneru, Tanneru, Sree Harsha, Chirag Agarwal +3 · 2 citations
Computer Science · #Topic Modeling #Explainable Artificial Intelligence (XAI) #Natural Language Processing Techniques
- Soft Best-of-n Sampling for Model Alignment
2025/05/06 by Claudio Mayrink Verdun, Alex Oesterling, Verdun, Claudio Mayrink +5 · 5 citations
Computer Science · Engineering · #Machine Learning and Algorithms #Machine Learning and Data Classification #Reservoir Engineering and Simulation Methods
- Robust and Stable Black Box Explanations
2020/11/12 by Himabindu Lakkaraju, Nino Arsov, Lakkaraju, Himabindu +3 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification
- More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
2024/04/29 by Aaron J. Li, Satyapriya Krishna, Li, Aaron J. +3 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques
- How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
2025/04/03 by Hongzhe Du, Du, Hongzhe, Weikai Li +13 · 2 citations
Computer Science · Medicine · #Explainable Artificial Intelligence (XAI) #Artificial Intelligence in Healthcare and Education #Topic Modeling
- Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems
2025/05/30 by Shichang Zhang, Zhang, Shichang, Hongguang Du +5 · 1 citation
Social Sciences · Computer Science · #Ethics and Social Impacts of AI #Blockchain Technology Applications and Security
- EvoLM: In Search of Lost Language Model Training Dynamics
2025/06/19 by Zhenting Qi, Qi, Zhenting, Fan Nie +15 · 1 voice · 1 citation
Computer Science · Medicine · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Healthcare and Education #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG