vix.ing · top · new · best · stats · spec

Extrapolation Guarantees for Perturbation Modeling Under the Additive Latent Shift Assumption

2025/04/25 by Julius von Kügelgen, Jakob Ketterer, von Kügelgen, Julius +9 · 1 citation
Biochemistry, Genetics and Molecular Biology · Computer Science · #Bayesian Methods and Mixture Models #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Single-cell and spatial transcriptomics

paper · pdf · doi:10.48550/arxiv.2504.18522

openalex publication_date 2025/04/25 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29

Abstract

We consider the problem of modeling the effects of perturbations like gene knockouts on measurements such as single-cell RNA counts. Given data for some perturbations, we aim to predict the distribution of measurements for new combinations of perturbations. To address this challenging extrapolation task, we posit that perturbations act additively in a suitable, unknown embedding space. We formulate the data-generating process as a latent variable model, in which perturbations amount to mean shifts in latent space and can be combined additively. We then prove that, given sufficiently diverse training perturbations, the representation and perturbation effects are identifiable up to orthogonal transformation and use this to derive extrapolation guarantees for unseen perturbations that can be expressed as linear combinations of seen ones. To estimate the model from data, we propose the perturbation distribution autoencoder (PDAE), which is trained by maximizing the distributional similarity between true and simulated perturbation distributions. The trained model can then be used to predict previously unseen perturbation distributions. In support of our theoretical results, we demonstrate through simulations that PDAE can accurately predict the effects of unseen but identifiable perturbations, and showcase the method on combinatorial gene perturbation data.

Cited by

Related