Local Haplotype Classifiers enable Efficient, Flexible, and Secure Genotype Imputation and Downstream Analyses
Cheema, M. N.; Nazir, A.; Moon, J.; Oh, Y.; Naseri, A.; Zhi, D.; Jiang, X.; Kim, M.; Harmanci, A. O.
Show abstract
The decreasing cost of genotyping technologies led to abundant availability and usage of genetic data. Although it offers many potentials for improving health and curing diseases, genetic data is highly intrusive in many aspects of individual privacy. Secure genotype analysis methods have been developed to perform numerous tasks such as genome-wide association studies, meta-analysis, kinship inference, and genotype imputation outsourcing. Here we present a new approach for using lightweight haplotype classifier models to use predicted haplotype information in a flexible privacy-preserving framework to perform genotype imputation and downstream tasks. Compared to the previous secure methods that rely main on linear models, our approach utilizes efficient models that rely on utilizing haplotypic information, which improves accuracy and increases the throughput of imputation by performing multiple imputations per model evaluation.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Assessing transcriptomic re-identification risks using discriminative sequence models 95%
- ML-MAGES: A machine learning framework for multivariate genetic association analyses with genes and effect size shrinkage 95%
- Secure Discovery of Genetic Relatives across Large-Scale and Distributed Genomic Datasets 95%
Similar papers in this journal
- Primo: integration of multiple GWAS and omics QTL summary statistics for elucidation of molecular mechanisms of trait-associated SNPs and detection of pleiotropy in complex traits 94%
- Optimizing and benchmarking polygenic risk scores with GWAS summary statistics 94%
- Fine-tuning sequence-to-expression models onpersonal genome and transcriptome data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.