Toward Causal Representation Learning
2021/02/26 by Bernhard Schölkopf, Bernhard Scholkopf, Francesco Locatello +5 · 1,061 citations
Computer Science · Psychology · #AI-based Problem Solving and Planning #Anomaly Detection Techniques and Applications #Artificial intelligence #Bayesian Modeling and Causal Inference #Cognitive psychology #Cognitive science #Computer science #Political science #Psychology #Representation (politics)
paper · pdf · doi:10.1109/jproc.2021.3058954
published in Proceedings of the IEEE 109(5), 612-634 (Institute of Electrical and Electronics Engineers)
openalex publication_date 2021/02/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Abstract
The two fields of machine learning and graphical causality arose and are developed separately. However, there is, now, cross-pollination and increasing interest in both fields to benefit from the advances of the other. In this article, we review fundamental concepts of causal inference and relate them to crucial open problems of machine learning, including transfer and generalization, thereby assaying how causality can contribute to modern machine learning research. This also applies in the opposite direction: we note that most work in causality starts from the premise that the causal variables are given. A central problem for AI and causality is, thus, causal representation learning, that is, the discovery of high-level causal variables from low-level observations. Finally, we delineate some implications of causality for machine learning and propose key research areas at the intersection of both communities.
Citations
Cited by
- Contextualizing predictive minds
- Domain generalization via causal fine-grained feature decomposition and learning
- On the Identifiability of Controlled World Models
- Respecting causality for training physics-informed neural networks
- Knowledge-Guided Time-Varying Causal Inference for Arctic Sea Ice Dynamics
- To Use AI as Dice of Possibilities with Timing Computation
- Interventional Causal Circuits for Safe Robot Action Testing and Failure Recovery
- Learning-Driven Channel Representation for Wireless Localization: From Channel Observations to Location Inference
- Teacher-Guided Causal Interventions for Image Denoising: Orthogonal Content-Noise Disentanglement in Vision Transformers
- Causal Supervision of Attention for Affective Behaviour Analysis
- Separable Pathways for Causal Reasoning: How Architectural Scaffolding Enables Hypothesis-Space Restructuring in LLM Agents
- Challenges in Statistics: A Dozen Challenges in Causality and Causal Inference
- The Cake that is Intelligence and Who Gets to Bake it: An AI Analogy and its Implications for Participation
- Unsupervised Object Learning via Common Fate
- Integrating Expert ODEs into Neural ODEs: Pharmacology and Disease Progression
- Artificial General Intelligence (AGI)-Native Wireless Systems: A Journey Beyond 6G
- Re-examining Granger Causality with Causal Bayesian Networks and Reichenbachs Principles
- Beyond ICA: Identifiability by Symmetry Breaking
- polyDAG: Polynomial Acyclicity Constraints for Efficient Continuous Causal Discovery in Visual Semantic Graphs
- Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models
- The Causal Loss: Driving Correlation to Imply Causation
- Invariant Feature Extraction Through Conditional Independence and the Optimal Transport Barycenter Problem: the Gaussian case
- The Epistemological Consequences of Large Language Models: Rethinking collective intelligence and institutional knowledge
- Toward Scalable and Valid Conditional Independence Testing with Spectral Representations
- Learning General Policies with Policy Gradient Methods
- Multi-Part Object Representations via Graph Structures and Co-Part Discovery
- Fairness Is Not Just Ethical: Performance Trade-Off via Data Correlation Tuning to Mitigate Bias in ML Software
- Unifying Causal Reinforcement Learning: Survey, Taxonomy, Algorithms and Applications
- Domain-Agnostic Causal-Aware Audio Transformer for Infant Cry Classification
- LINA: Learning INterventions Adaptively for Physical Alignment and Generalization in Diffusion Models
- CausalCLIP: Causally-Informed Feature Disentanglement and Filtering for Generalizable Detection of Generated Images
- Compressed Causal Reasoning: Quantization and GraphRAG Effects on Interventional and Counterfactual Accuracy
- Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
- Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality
- Causal Attribution of Model Performance Gaps in Medical Imaging Under Distribution Shifts
- Modular Jets for Supervised Pipelines: Diagnosing Mirage vs Identifiability
- Learning Causal Semantic Representation for Out-of-Distribution Prediction
- Vision and Causal Learning Based Channel Estimation for THz Communications
- Stress-Testing Causal Claims via Cardinality Repairs
- Disentanglement via Mechanism Sparsity Regularization: A New Principle for Nonlinear ICA
- PISA: Prioritized Invariant Subgraph Aggregation
- On the Transfer of Disentangled Representations in Realistic Settings
- CID: Measuring Feature Importance Through Counterfactual Distributions
- Typing assumptions improve identification in causal discovery
- Relaxing partition admissibility in Cluster-DAGs: a causal calculus with arbitrary variable clustering
- Mitigating Length Bias in RLHF through a Causal Lens
- Dynamic Sparsity: Challenging Common Sparsity Assumptions for Learning World Models in Robotic Reinforcement Learning Benchmarks
- CaberNet: Causal Representation Learning for Cross-Domain HVAC Energy Prediction
- Causal Structure and Representation Learning with Biomedical Applications
- Towards Causal Market Simulators
- Causal Graph Neural Networks for Healthcare
- DoFlow: Flow-based Generative Models for Interventional and Counterfactual Forecasting on Time Series
- Understanding Hardness of Vision-Language Compositionality from A Token-level Causal Lens
- Causal thinking for decision making on Electronic Health Records: why and how
- Semantics-Native Communication with Contextual Reasoning
- Generalizing Graph Neural Networks on Out-Of-Distribution Graphs
- G2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal Disentanglement
- Safety from Honesty in a Disinterested AI Predictor
- Causal Disentangled Recommendation against User Preference Shifts
- Use and usability: concepts of representation in philosophy, neuroscience, cognitive science, and computer science
- FELA: A Multi-Agent Evolutionary System for Feature Engineering of Industrial Event Log Data
- Eigenfunction Extraction for Ordered Representation Learning
- Beyond Prompt Engineering: Neuro-Symbolic-Causal Architecture for Robust Multi-Objective AI Agents
- Empowering Multimodal Respiratory Sound Classification with Counterfactual Adversarial Debiasing for Out-of-Distribution Robustness
- Resilient Radio Access Networks: AI and the Unknown Unknowns
- ROPES: Robotic Pose Estimation via Score-Based Causal Representation Learning
- LyTimeT: Towards Robust and Interpretable State-Variable Discovery
- Towards Single-Source Domain Generalized Object Detection via Causal Visual Prompts
- Online Time Series Forecasting with Theoretical Guarantees
- Diverse Influence Component Analysis: A Geometric Approach to Nonlinear Mixture Identifiability
- A Preliminary Exploration of the Differences and Conjunction of Traditional PNT and Brain-inspired PNT
- Foundation Models for Scientific Discovery: From Paradigm Enhancement to Paradigm Transition
- Causal Time Series Modeling of Supraglacial Lake Evolution in Greenland under Distribution Shift
- Doubly Robust Estimation of Causal Effects in Strategic Equilibrium Systems
- Measure-Theoretic Anti-Causal Representation Learning
- Exploratory Causal Inference in SAEnce
- CausalVerse: Benchmarking Causal Representation Learning with Configurable High-Fidelity Simulations
- Generalization and Robustness Implications in Object-Centric Learning
- Sculpting Latent Spaces With MMD: Disentanglement With Programmable Priors
- Causal Disentanglement Learning for Accurate Anomaly Detection in Multivariate Time Series
- Explainable Human-in-the-Loop Segmentation via Critic Feedback Signals
- Counterfactual Identifiability via Dynamic Optimal Transport
- Causality Guided Representation Learning for Cross-Style Hate Speech Detection
- Executable Counterfactuals: Improving LLMs' Causal Reasoning Through Code
- Root Cause Analysis of Outliers in Unknown Cyclic Graphs
- Provable Affine Identifiability of Nonlinear CCA under Latent Distributional Priors
- From Pixels to Factors: Learning Independently Controllable State Variables for Reinforcement Learning
- Distributionally Robust Causal Abstractions
- Nonparametric Identification of Latent Concepts
- Causal-Adapter: Taming Text-to-Image Diffusion for Faithful Counterfactual Generation
- LLM Interpretability with Identifiable Temporal-Instantaneous Representation
- From explainable to interpretable deep learning for natural language processing in healthcare: How far from reality?
- Partially Observed Structural Causal Models
- Understanding Catastrophic Interference: On the Identifibility of Latent Representations
- Learning Neural Causal Models with Active Interventions
- Linear Causal Representation Learning by Topological Ordering, Pruning, and Disentanglement
- Mechanistic Independence: A Principle for Identifiable Disentangled Representations
- Practical do-Shapley Explanations with Estimand-Agnostic Causal Inference
- A perspective of advances in optical methods for biological sample characterization
- Context-Informed Ship Trajectory Prediction via Conditional Attention
- From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery
- Seeing the unseen: A novel approach to extract latent plant root traits from digital images
- Preventing Spurious Interactions: A New Inductive Bias for Accurate Treatment Effect Estimation
- Predicting cognitive function 3 months after surgery in patients with a glioma
- Interpretation, extrapolation and perturbation of single cells
- Towards Causal Representation Learning with Observable Sources as Auxiliaries
- Revealing Multimodal Causality with Large Language Models
- Data Complexity: a threshold between Classical and Quantum Machine Learning -- Part I
- Consistent Bayesian causal discovery for structural equation models with equal error variances
- Deceptive Risk Minimization: Out-of-Distribution Generalization by Deceiving Distribution Shift Detectors
- PersonaX: Multimodal Datasets with LLM-Inferred Behavior Traits
- Abduct, Act, Predict: Scaffolding Causal Inference for Automated Failure Attribution in Multi-Agent Systems
- Boxhead: A Dataset for Learning Hierarchical Representations
- Learning Object-Centric Representations in SAR Images with Multi-Level Feature Fusion
- Directed Cyclic Graph for Causal Discovery from Multivariate Functional Data
- Towards explainable decision support using hybrid neural models for logistic terminal automation
- Learning latent causal graphs via mixture oracles
- Effects of Distributional Biases on Gradient-Based Causal Discovery in the Bivariate Categorical Case
- Towards Out-Of-Distribution Generalization: A Survey
- ChainReaction: Causal Chain-Guided Reasoning for Modular and Explainable Causal-Why Video Question Answering
- When Is Causal Inference Possible? A Statistical Test for Unmeasured Confounding
- On Disentangled Representations Learned From Correlated Data
- Self-Supervised Learning with Data Augmentations Provably Isolates Content from Style
- Conterfactual Generative Zero-Shot Semantic Segmentation
- Locality-aware Concept Bottleneck Model
- LOOP: A Plug-and-Play Neuro-Symbolic Framework for Enhancing Planning in Autonomous Systems
- A Topological Perspective on Causal Inference
- Root Cause Analysis of Hydrogen Bond Separation in Spatio-Temporal Molecular Dynamics using Causal Models
- Scientific Machine Learning Through Physics–Informed Neural Networks: Where we are and What’s Next
- Structured Kernel Regression VAE: A Computationally Efficient Surrogate for GP-VAEs in ICA
- Information Bottleneck-based Causal Attention for Multi-label Medical Image Recognition
- Algorithmic Fairness amid Social Determinants: Reflection, Characterization, and Approach
- Learning Causal Structure Distributions for Robust Planning
- Learning Robust Intervention Representations with Delta Embeddings
- CAMA: Enhancing Mathematical Reasoning in Large Language Models with Causal Knowledge
- Beyond Simulation: Benchmarking World Models for Planning and Causality in Autonomous Driving
- Compositional Video Synthesis by Temporal Object-Centric Learning
- Trek-Based Parameter Identification for Linear Causal Models With Arbitrarily Structured Latent Variables
- Artificial Intelligence and Democracy: A Conceptual Framework
- Formal Modeling as Theoretical Glue Between Laboratory and Naturalistic Studies of Memory
- Causal Mechanism Estimation in Multi-Sensor Systems Across Multiple Domains
- Should Bias Always be Eliminated? A Principled Framework to Use Data Bias for OOD Generation
- Causal Process Models: Reframing Dynamic Causal Graph Discovery as a Reinforcement Learning Problem
- Optimal Empirical Risk Minimization under Temporal Distribution Shifts
- Reimagining an autonomous vehicle
- Optimization-based Causal Estimation from Heterogenous Environments
- Conspiracy to Commit: Information Pollution, Artificial Intelligence, and Real-World Hate Crime
- Searching for actual causes: Approximate algorithms with adjustable precision
- What Has a Foundation Model Found? Using Inductive Bias to Probe for World Models
- Cross-Modal Dual-Causal Learning for Long-Term Action Recognition
- Towards Principled Disentanglement for Domain Generalization
- Causal Foundation Models: Disentangling Physics from Instrument Properties
- Data Supplement to the paper "Intervening to Learn and Compose Causally Disentangled Representations"
- Incorporating Interventional Independence Improves Robustness against Interventional Distribution Shift
- When Data-Free Knowledge Distillation Meets Non-Transferable Teacher: Escaping Out-of-Distribution Trap is All You Need
- Uncertainty-Aware Deepfake Detection via Multi-View Structural Learning
- L-VAE: Variational Auto-Encoder with Learnable Beta for Disentangled Representation
- A Survey on Causal Discovery: Theory and Practice
- Causal Discovery of Latent Variables in Galactic Archaeology
- A New Perspective On AI Safety Through Control Theory Methodologies
- Curious Causality-Seeking Agents Learn Meta Causal World
- Bayesian Invariance Modeling of Multi-Environment Data
- Active Inference AI Systems for Scientific Discovery
- Exploring Graph-Transformer Out-of-Distribution Generalization Abilities
- Domain Knowledge in Artificial Intelligence: Using Conceptual Modeling to Increase Machine Learning Accuracy and Explainability
- Causal Representation Learning with Observational Grouping for CXR Classification
- Tagged for Direction: Pinning Down Causal Edge Directions with Precision
- On using AI for EEG-based BCI applications: problems, current challenges and future trends
- Dynamic Inference with Neural Interpreters
- Sequential Causal Discovery with Noisy Language Model Priors
- Causally Steered Diffusion for Automated Video Counterfactual Generation
- Interpretable Causal Representation Learning for Biological Data in the Pathway Space
- Causality in the human niche: lessons for machine learning
- DISCO: Mitigating Bias in Deep Learning with Conditional Distance Correlation
- The development of human causal learning and reasoning
- Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
- Causal Climate Emulation with Bayesian Filtering
- Efficiently Disentangle Causal Representations
- Preference Learning for AI Alignment: a Causal Perspective
- Post-discovery Analysis of Anomalous Subsets
- Text Data Augmentation for Deep Learning
- Learning Treatment Representations for Downstream Instrumental Variable Regression
- Bias as a Virtue: Rethinking Generalization under Distribution Shifts
- Understanding the European energy crisis through structural causal models
- From Invariant Representations to Invariant Data: Provable Robustness to Spurious Correlations via Noisy Counterfactual Matching
- A Reverse Causal Framework to Mitigate Spurious Correlations for Debiasing Scene Graph Generation
- Multi-bearing fault diagnosis method based on convolutional autoencoder causal decoupling domain generalization
- Exploring Novel Pooling Strategies for Edge Preserved Feature Maps in Convolutional Neural Networks
- Autoencoding Random Forests
- A Theoretical Analysis of Compositional Generalization in Neural Networks: A Necessary and Sufficient Condition
- Mitigating Context Bias in Domain Adaptation for Object Detection using Mask Pooling
- CockpitHAT: Dependency-Graph-Driven Hierarchical Attribution for Embodied Multi-Agent Cockpits
- The Third Pillar of Causal Analysis? A Measurement Perspective on Causal Representations
- Towards Identifiability of Interventional Stochastic Differential Equations
- CausalVLBench: Benchmarking Visual Causal Reasoning in Large Vision-Language Models
- Virtual Cells: Predict, Explain, Discover
- Interpretable Neural System Dynamics: Combining Deep Learning with System Dynamics Modeling to Support Critical Applications
- Causal Cartographer: From Mapping to Reasoning Over Counterfactual Worlds
- Causality-Inspired Robustness for Nonlinear Models via Representation Learning
- Quantum Algorithms for Causal Estimands
- Independent mechanism analysis, a new concept?
- Information Science Principles of Machine Learning: A Causal Chain Meta-Framework Based on Formalized Information Mapping
- Denoising Mutual Knowledge Distillation in Bi-Directional Multiple Instance Learning
- Artificial intelligence and illusions of understanding in scientific research
- Causal discovery on vector-valued variables and consistency-guided aggregation
- On Measuring Intrinsic Causal Attributions in Deep Neural Networks
- Parameter Estimation using Reinforcement Learning Causal Curiosity: Limits and Challenges
- Modeling Unseen Environments with Language-guided Composable Causal Components in Reinforcement Learning
- Towards Unbiased Visual Emotion Recognition via Causal Intervention
- Causal View of Time Series Imputation: Some Identification Results on Missing Mechanism
- Directed Acyclic Graph Learning on Attributed Heterogeneous Network
- Contextures: Representations from Contexts
- Generative Engine Optimization: A VLM and Agent Framework for Pinterest Acquisition Growth
- Querying Counterfactuals on Tissue Graphs with Supervised Disentanglement
- A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
- When Does LeJEPA Learn a World Model?
- If Concept Bottlenecks are the Question, are Foundation Models the Answer?
- AI Alignment in Medical Imaging: Unveiling Hidden Biases Through Counterfactual Analysis
- Extrapolation Guarantees for Perturbation Modeling Under the Additive Latent Shift Assumption
- Enhancing System Self-Awareness and Trust of AI: A Case Study in Trajectory Prediction and Planning
- CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention
- SVI-DAG: A Structured Variational Inference Approach to Bayesian Causal Discovery
- The RAIL Principles for Neurosymbolic AI: Reasoning, Assurances, Interfacing and Learning
- Disentangling Long and Short-Term Interests for Recommendation
- Causal DAG Summarization (Full Version)
- Support Operation Factorization: Compositional Readout of Frozen Vision Encoders under Controlled Interventions
- Towards Interpretable Deep Generative Models via Causal Representation Learning
- A conceptual synthesis of causal assumptions for causal discovery and inference
- On the Value of Cross-Modal Misalignment in Multimodal Representation Learning
- Are We Done with Object-Centric Learning?
- Structure-guided graph attention learning for traffic conflict prediction under heterogeneous expressway geometries
- Interpretable machine learning for genomics. [europepmc]
- Diffused responsibility: attributions of responsibility in the use of AI-driven clinical decision support systems. [europepmc]
- Can Robots Do Epidemiology? Machine Learning, Causal Inference, and Predicting the Outcomes of Public Health Interventions. [europepmc]
- Causal machine learning for healthcare and precision medicine. [europepmc]
- Machine Learning for Causal Inference in Biological Networks: Perspectives of This Challenge. [europepmc]
- End-to-end sequence-structure-function meta-learning predicts genome-wide chemical-protein interactions for dark proteins. [europepmc]
- Neural Information Squeezer for Causal Emergence. [europepmc]
- Generalising uncertainty improves accuracy and safety of deep learning analytics applied to oncology. [europepmc]
- The role of explainable Artificial Intelligence in high-stakes decision-making systems: a systematic review. [europepmc]
- Evaluating vaccine allocation strategies using simulation-assisted causal modeling. [europepmc]
- HEAR4Health: a blueprint for making computer audition a staple of modern healthcare. [europepmc]
- Artificial intelligence (AI)-it's the end of the tox as we know it (and I feel fine). [europepmc]
- Plant science in the age of simulation intelligence. [europepmc]
- Learning representations for image-based profiling of perturbations. [europepmc]
- Emergence and Causality in Complex Systems: A Survey of Causal Emergence and Related Quantitative Studies. [europepmc]
- Computer vision for plant pathology: A review with examples from cocoa agriculture. [europepmc]
- The transformative potential of artificial intelligence in solid organ transplantation. [europepmc]
- Eight challenges in developing theory of intelligence. [europepmc]
- Knowledge-based inductive bias and domain adaptation for cell type annotation. [europepmc]
- Step-by-step causal analysis of EHRs to ground decision-making. [europepmc]
- A large-scale benchmark for network inference from single-cell perturbation data. [europepmc]
- Early warning of complex climate risk with integrated artificial intelligence. [europepmc]
- Causality, Machine Learning, and Feature Selection: A Survey. [europepmc]
- Causally-Informed Instance-Wise Feature Selection for Explaining Visual Classifiers. [europepmc]
- Combinatorial prediction of therapeutic perturbations using causally inspired neural networks. [europepmc]
Related