English

From Identifiable Causal Representations to Controllable Counterfactual Generation: A Survey on Causal Generative Modeling

Machine Learning 2024-05-24 v2 Artificial Intelligence Machine Learning

Abstract

Deep generative models have shown tremendous capability in data density estimation and data generation from finite samples. While these models have shown impressive performance by learning correlations among features in the data, some fundamental shortcomings are their lack of explainability, tendency to induce spurious correlations, and poor out-of-distribution extrapolation. To remedy such challenges, recent work has proposed a shift toward causal generative models. Causal models offer several beneficial properties to deep generative models, such as distribution shift robustness, fairness, and interpretability. Structural causal models (SCMs) describe data-generating processes and model complex causal relationships and mechanisms among variables in a system. Thus, SCMs can naturally be combined with deep generative models. We provide a technical survey on causal generative modeling categorized into causal representation learning and controllable counterfactual generation methods. We focus on fundamental theory, methodology, drawbacks, datasets, and metrics. Then, we cover applications of causal generative models in fairness, privacy, out-of-distribution generalization, precision medicine, and biological sciences. Lastly, we discuss open problems and fruitful research directions for future work in the field.

Keywords

Cite

@article{arxiv.2310.11011,
  title  = {From Identifiable Causal Representations to Controllable Counterfactual Generation: A Survey on Causal Generative Modeling},
  author = {Aneesh Komanduri and Xintao Wu and Yongkai Wu and Feng Chen},
  journal= {arXiv preprint arXiv:2310.11011},
  year   = {2024}
}

Comments

Published in Transactions on Machine Learning Research (TMLR) (05/2024); 72 pages, 27 figures, 4 tables