PopGenAdapt: Semi-Supervised Domain Adaptation for Genotype-to-Phenotype Prediction in Underrepresented Populations
Comajoan Cara, M.; Mas Montserrat, D.; Ioannidis, A. G.
Show abstract
The lack of diversity in genomic datasets, currently skewed towards individuals of European ancestry, presents a challenge in developing inclusive biomedical models. The scarcity of such data is particularly evident in labeled datasets that include genomic data linked to electronic health records. To address this gap, this paper presents PopGenAdapt, a genotype-to-phenotype prediction model which adopts semi-supervised domain adaptation (SSDA) techniques originally proposed for computer vision. PopGenAdapt is designed to leverage the substantial labeled data available from individuals of European ancestry, as well as the limited labeled and the larger amount of unlabeled data from currently underrepresented populations. The method is evaluated in underrepresented populations from Nigeria, Sri Lanka, and Hawaii for the prediction of several disease outcomes. The results suggest a significant improvement in the performance of genotype-to-phenotype models for these populations over state-of-the-art supervised learning methods, setting SSDA as a promising strategy for creating more inclusive machine learning models in biomedical research. Our code is available at https://github.com/AI-sandbox/PopGenAdapt.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- CaLMPhosKAN: Prediction of General Phosphorylation Sites in Proteins via Fusion of Codon Aware Embeddings with Amino Acid Aware Embeddings and Wavelet-based Kolmogorov Arnold Network 94%
- Multi-Omic Graph Diagnosis (MOGDx) : A data integration tool to perform classification tasks for heterogeneous diseases 94%
- High-dimensional Biomarker Identification for Scalable and Interpretable Disease Prediction via Machine Learning Models 93%
Similar papers in this journal
- Computationally scalable regression modeling for ultrahigh-dimensional omics data with ParProx 94%
- Novel multi-omics deconfounding variational autoencoders can obtain meaningful disease subtyping 94%
- Synthetic observations from deep generative models and binary omics data with limited sample size 94%
Similar papers in this journal
Similar papers in this journal
- Federated Learning for multi-omics: a performance evaluation in Parkinson's disease 95%
- RiskPath : Explainable deep learning for multistep biomedical prediction in longitudinal data 94%
- Generating hard-to-obtain information from easy-to-obtain information: applications in drug discovery and clinical inference 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.