English

SNPgen: Phenotype-Supervised Genotype Representation and Synthetic Data Generation via Latent Diffusion

Machine Learning 2026-03-12 v1 Genomics

Abstract

Polygenic risk scores and other genomic analyses require large individual-level genotype datasets, yet strict data access restrictions impede sharing. Synthetic genotype generation offers a privacy-preserving alternative, but most existing methods operate unconditionally, producing samples without phenotype alignment, or rely on unsupervised compression, creating a gap between statistical fidelity and downstream task utility. We present SNPgen, a two-stage conditional latent diffusion framework for generating phenotype-supervised synthetic genotypes. SNPgen combines GWAS-guided variant selection (1,024-2,048 trait-associated SNPs) with a variational autoencoder for genotype compression and a latent diffusion model conditioned on binary disease labels via classifier-free guidance. Evaluated on 458,724 UK Biobank individuals across four complex diseases (coronary artery disease, breast cancer, type 1 and type 2 diabetes), models trained on synthetic data matched real-data predictive performance in a train-on-synthetic, test-on-real protocol, approaching genome-wide PRS methods that use 22-6×6\times more variants. Privacy analysis confirmed zero identical matches, near-random membership inference (AUC 0.50\approx 0.50), preserved linkage disequilibrium structure, and high allele frequency correlation (r0.95r \geq 0.95) with source data. A controlled simulation with known causal effects verified faithful recovery of the imposed genetic association structure.

Keywords

Cite

@article{arxiv.2603.10873,
  title  = {SNPgen: Phenotype-Supervised Genotype Representation and Synthetic Data Generation via Latent Diffusion},
  author = {Andrea Lampis and Michela Carlotta Massi and Nicola Pirastu and Francesca Ieva and Matteo Matteucci and Emanuele Di Angelantonio},
  journal= {arXiv preprint arXiv:2603.10873},
  year   = {2026}
}
R2 v1 2026-07-01T11:14:49.676Z