Trustworthy AI: From Principles to Practices
2022/08/18 by Bo Li, Peng Qi, Bo Liu +5 · 61 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Privacy-Preserving Technologies in Data
paper · pdf · doi:10.1145/3555803
openalex publication_date 2022/08/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
Abstract
The rapid development of Artificial Intelligence (AI) technology has enabled the deployment of various systems based on it. However, many current AI systems are found vulnerable to imperceptible attacks, biased against underrepresented groups, lacking in user privacy protection. These shortcomings degrade user experience and erode people’s trust in all AI systems. In this review, we provide AI practitioners with a comprehensive guide for building trustworthy AI systems. We first introduce the theoretical framework of important aspects of AI trustworthiness, including robustness, generalization, explainability, transparency, reproducibility, fairness, privacy preservation, and accountability. To unify currently available but fragmented approaches toward trustworthy AI, we organize them in a systematic approach that considers the entire lifecycle of AI systems, ranging from data acquisition to model development, to system development and deployment, finally to continuous monitoring and governance. In this framework, we offer concrete action items for practitioners and societal stakeholders (e.g., researchers, engineers, and regulators) to improve AI trustworthiness. Finally, we identify key opportunities and challenges for the future development of trustworthy AI systems, where we identify the need for a paradigm shift toward comprehensively trustworthy AI systems.
Citations
Cited by
- Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response
- Physical AI Governance: From Theory to Practice Across Life Cycle
- Can Large Language Models Function as Qualified Pediatricians? A Systematic Evaluation in Real-World Clinical Contexts
- Stabilizing Multi-Attack Adversarial Training via Bandit Optimization
- "Show Me You Comply... Without Showing Me Anything": Zero-Knowledge Software Auditing for AI-Enabled Systems
- Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels
- Revolutionizing healthcare data analytics with federated learning: A comprehensive survey of applications, systems, and future directions
- HIKMA: Human-Inspired Knowledge by Machine Agents through a Multi-Agent Framework for Semi-Autonomous Scientific Conferences
- Weak-to-Strong Generalization under Distribution Shifts
- Understanding AI Trustworthiness: A Scoping Review of AIES & FAccT Articles
- Opening up ChatGPT: Tracking openness, transparency, and accountability in instruction-tuned text generators
- An Empirical Study on Variance-based MC Dropout Uncertainty-Error Correlation in 2D Brain Tumor Segmentation
- LLMartini: Seamless and Interactive Leveraging of Multiple LLMs through Comparison and Composition
- Towards Robust Artificial Intelligence: Self-Supervised Learning Approach for Out-of-Distribution Detection
- Are Large Language Models Effective Knowledge Graph Constructors?
- FML-bench: A Benchmark for Automatic ML Research Agents Highlighting the Importance of Exploration Breadth
- Private and Fair Machine Learning: Revisiting the Disparate Impact of Differentially Private SGD
- Walk the Talk: Is Your Log-based Software Reliability Maintenance System Really Reliable?
- Interpretable deep learning: interpretation, interpretability, trustworthiness, and beyond
- Generation-Time vs. Post-hoc Citation: A Holistic Evaluation of LLM Attribution
- Challenges and strategies for wide-scale artificial intelligence (AI) deployment in healthcare practices: A perspective for healthcare organizations
- Re-evaluating GPT-4’s bar exam performance
- How to teach responsible AI in Higher Education: challenges and opportunities
- Perspectives and potential issues in using artificial intelligence for computer science education
- Twenty-four years of empirical research on trust in AI: a bibliometric review of trends, overlooked issues, and future directions
- Neuro-Symbolic Frameworks: Conceptual Characterization and Empirical Comparative Analysis
- Benchmarking Vision Transformers and CNNs for Thermal Photovoltaic Fault Detection with Explainable AI Validation
- An Investigation of Visual Foundation Models Robustness
- Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations
- TechOps: Technical Documentation Templates for the AI Act
- Conformal Prediction and Trustworthy AI
- Dimensional Characterization and Pathway Modeling for Catastrophic AI Risks
- Evolution of AI Agent Registry Solutions: Centralized, Enterprise, and Distributed Approaches
- A Survey on Data Security in Large Language Models
- AutoSIGHT: Automatic Eye Tracking-based System for Immediate Grading of Human experTise
- Challenges of Trustworthy Federated Learning: What's Done, Current Trends and Remaining Work
- Can AI Model the Complexities of Human Moral Decision-making? A Qualitative Study of Kidney Allocation Decisions
- TAI Scan Tool: A RAG-Based Tool With Minimalistic Input for Trustworthy AI Self-Assessment
- MUPAX: Multidimensional Problem Agnostic eXplainable AI
- ZKPROV: A Zero-Knowledge Approach to Dataset Provenance for Large Language Models
- The Singapore Consensus on Global AI Safety Research Priorities
- Accountability of Robust and Reliable AI-Enabled Systems: A Preliminary Study and Roadmap
- TRUST: Transparent, Robust and Ultra-Sparse Trees
- Sampling Preferences Yields Simple Trustworthiness Scores
- A Trustworthiness-based Metaphysics of Artificial Intelligence Systems
- Improving Reliability and Explainability of Medical Question Answering through Atomic Fact Checking in Retrieval-Augmented LLMs
- Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies
- Engineering Trustworthy Machine-Learning Operations with Zero-Knowledge Proofs
- Towards Generalized Proactive Defense against Face Swapping with Contour-Hybrid Watermark
- Trustworthy AI in Digital Health: A Comprehensive Review of Robustness and Explainability
- Aligning Trustworthy AI with Democracy: A Dual Taxonomy of Opportunities and Risks
- Toward Adaptive Categories: Dimensional Governance for Agentic AI
- Counterfactual Reasoning for Causal Responsibility Attribution in Probabilistic Multi-Agent Systems
- Edge-Cloud Collaborative Computing on Distributed Intelligence and Model Optimization: A Survey
- FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics
- Physical foundations for trustworthy medical imaging: a review for artificial intelligence researchers
- On the Need to Rethink Trust in AI Assistants for Software Development: A Critical Review
- Building Trustworthy Multimodal AI: A Review of Fairness, Transparency, and Ethics in Vision-Language Tasks
- Artificial Intelligence, Structure of Knowledge, and the Future Directions for Macromarketing
- Trustworthy AI Must Account for Interactions
- Explaining Uncertainty in Multiple Sclerosis Lesion Segmentation Beyond Prediction Errors
Related