Prediction Of Eye, Hair And Skin Color In Admixed Populations Of Latin America
Palmal, S.; Adhikari, K.; Mendoza-Revilla, J.; Fuentes-Guajardo, M.; Silva de Cerqueira, C. C.; Chacon-Duque, J. C.; Sohail, A.; Hurtado, M.; Villegas, V.; Granja, V.; Jaramillo, C.; Arias, W.; Barquera Lozano, R.; Everardo Martinez, P.; Gomez-Valdes, J.; Villamil-Ramirez, H.; Hunemeier, T.; Ramallo, V.; Gonzalez-Jose, R.; Schuler-Faccini, L.; Bortolini, M.-C.; Acuna-Alonzo, V.; Canizales-Quinteros, S.; Gallo, C.; Poletti, G.; Bedoya, G.; Rothhammer, F.; Balding, D.; Faux, P.; Ruiz-Linares, A.
Show abstract
We report an evaluation of prediction accuracy for eye, hair and skin pigmentation based on genomic and phenotypic data for over 6,500 admixed Latin Americans (the CANDELA dataset). We examined the impact on prediction accuracy of three main factors: (i) The methods of prediction, including classical statistical methods and machine learning approaches, (ii) The inclusion of non-genetic predictors, continental genetic ancestry and pigmentation SNPs in the prediction models, and (iii) Compared two sets of pigmentation SNPs: the commonly-used HIrisPlex-S set (developed in Europeans) and novel SNP sets we defined here based on genome-wide association results in the CANDELA sample. We find that Random Forest or regression are globally the best performing methods. Although continental genetic ancestry has substantial power for prediction of pigmentation in Latin Americans, the inclusion of pigmentation SNPs increases prediction accuracy considerably, particularly for skin color. For hair and eye color, HIrisPlex-S has a similar performance to the CANDELA-specific prediction SNP sets. However, for skin pigmentation the performance of HIrisPlex-S is markedly lower than the SNP set defined here, including predictions in an independent dataset of Native American data. These results reflect the relatively high variation in hair and eye color among Europeans for whom HIrisPlex-S was developed, whereas their variation in skin pigmentation is comparatively lower. Furthermore, we show that the dataset used in the training of prediction models strongly impacts on the portability of these models across Europeans and Native Americans.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The skin we live in: pigmentation traits and tanning behaviour in British young adults, an observational and genetically-informed study 91%
- A novel framework for analysis of the shared genetic background of correlated traits 91%
- The FORCE panel: An all-in-one SNP marker set for confirming investigative genetic genealogy leads and for general forensic applications 91%
Similar papers in this journal
- JointPRS: A Data-Adaptive Framework for Multi-Population Genetic Risk Prediction Incorporating Genetic Correlation 92%
- A novel method for an unbiased estimate of cross-ancestry genetic correlation using individual-level data 92%
- Probabilistic inference of the genetic architecture underlying functional enrichment of complex traits 91%
Similar papers in this journal
- Leveraging Global Genetics Resources to Enhance Polygenic Prediction Across Ancestrally Diverse Populations 91%
- A reference panel for linkage disequilibrium and genotype imputation using whole-genome sequencing data from 2,680 participants across India 91%
- Evaluating Genomic Polygenic Risk Scores for Childhood Acute Lymphoblastic Leukemia in Latinos 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.