ClOneHORT: Approaches for Improved Fidelity in Generative Models of Synthetic Genomes
Laboulaye, R.; Borda, V.; Chen, S.; North, K. E.; Kaplan, R.; O'Connor, T. D.
Show abstract
MotivationDeep generative models have the potential to overcome difficulties in sharing individual-level genomic data by producing synthetic genomes that preserve the genomic associations specific to a cohort while not violating the privacy of any individual cohort member. However, there is significant room for improvement in the fidelity and usability of existing synthetic genome approaches. ResultsWe demonstrate that when combined with plentiful data and with population-specific selection criteria, deep generative models can produce synthetic genomes and cohorts that closely model the original populations. Our methods improve fidelity in the site-frequency spectra and linkage disequilibrium decay and yield synthetic genomes that can be substituted in downstream local ancestry inference analysis, recreating results with .91 to .94 accuracy. AvailabilityThe model described in this paper is freely available at github.com/rlaboulaye/clonehort.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Large-scale Inference of Population Structure in Presence of Missingness using PCA 96%
- Computing Linkage Disequilibrium Aware Genome Embeddings using Autoencoders 96%
- ACTIVA: realistic single-cell RNA-seq generation with automatic cell-type identification using introspective variational autoencoders 95%
Similar papers in this journal
Similar papers in this journal
- Using feedback in pooled experiments augmented with imputation for high genotyping accuracy at reduced cost 93%
- CHARON: Estimating the drift time and the number of individuals in environmental DNA with diploid individuals 92%
- The promise and challenge of spatial inference with the full ancestral recombination graph underBrownian motion 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.