vix.ing · top · new · best · stats · spec

A theory of independent mechanisms for extrapolation in generative\n models

2020/03/31 by Michel Besserve, Rémy Sun, Besserve, Michel +5 · 1 citation
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification

paper · pdf · doi:10.48550/arxiv.2004.00184

openalex publication_date 2020/03/31 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28

Abstract

Generative models can be trained to emulate complex empirical data, but are\nthey useful to make predictions in the context of previously unobserved\nenvironments? An intuitive idea to promote such extrapolation capabilities is\nto have the architecture of such model reflect a causal graph of the true data\ngenerating process, such that one can intervene on each node independently of\nthe others. However, the nodes of this graph are usually unobserved, leading to\noverparameterization and lack of identifiability of the causal structure. We\ndevelop a theoretical framework to address this challenging situation by\ndefining a weaker form of identifiability, based on the principle of\nindependence of mechanisms. We demonstrate on toy examples that classical\nstochastic gradient descent can hinder the model's extrapolation capabilities,\nsuggesting independence of mechanisms should be enforced explicitly during\ntraining. Experiments on deep generative models trained on real world data\nsupport these insights and illustrate how the extrapolation capabilities of\nsuch models can be leveraged.\n

Cited by

Related