Predicting the genetic ancestry of 2.6 million New York City patients using clinical data
Ramlall, V.; Quinnies, K. M.; Vanguri, R.; Lorberbaum, T.; Goldstein, D. B.; Tatonetti, N. P.
Show abstract
Ancestry is an essential covariate in clinical genomics research. When genetic data are available, dimensionality reduction techniques, such as principal components analysis, are used to determine ancestry in complex populations. Unfortunately, these data are not always available in the clinical and research settings. For example, electronic health records (EHRs), which are a rich source of temporal human disease data that could be used to enhance genetic studies, do not directly capture ancestry. Here, we present a novel algorithm for predicting genetic ancestry using only variables that are routinely captured in EHRs, such as self-reported race and ethnicity, and condition billing codes. Using patients that have both genetic and clinical information at Columbia University/ New York-Presbyterian Irving Medical Center, we developed a pipeline that uses only clinical data to predict the genetic ancestry of all patients of which more than 80% identify as other or unknown. Our ancestry estimates can be used in observational studies of disease inheritance, to guide genetic cohort studies, or to explore health disparities in clinical care and outcomes.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 94%
- Genome-Wide Sequencing as a First-Tier Screening Test for Short Tandem Repeat Expansions 94%
- Evaluating Genome Sequencing Strategies: Trio, Singleton, and Standard Testing in Rare Disease Diagnosis 93%
Similar papers in this journal
- Long-read genome sequencing for the diagnosis of neurodevelopmental disorders 94%
- Inclusion of Variants Discovered from Diverse Populations Improves Polygenic Risk Score Transferability 94%
- Leveraging TOPMed Imputation Server and Constructing a Cohort-Specific Imputation Reference Panel to Enhance Genotype Imputation among Cystic Fibrosis Patients 94%
Similar papers in this journal
- Optimization of Multi-Ancestry Polygenic Risk Score Disease Prediction Models 94%
- Can imputation in a European country be improved by local reference panels? The example of France 93%
- The application of Large Language Models to the phenotype-based prioritization of causative genes in rare disease patients 93%
Similar papers in this journal
- The Importance of Automation in Genetic Diagnosis: Lessons from Analyzing an Inherited Retinal Degeneration Cohort with the Mendelian Analysis Toolkit (MATK) 95%
- Evaluating the impact of modeling choices on the performance of integrated genetic and clinical models 94%
- A gene pathogenicity tool 'GenePy' identifies missed biallelic diagnoses in the 100,000 Genomes Project 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.