Back

Predicting the genetic ancestry of 2.6 million New York City patients using clinical data

Ramlall, V.; Quinnies, K. M.; Vanguri, R.; Lorberbaum, T.; Goldstein, D. B.; Tatonetti, N. P.

2019-09-14 bioinformatics
10.1101/768440 bioRxiv
Show abstract

Ancestry is an essential covariate in clinical genomics research. When genetic data are available, dimensionality reduction techniques, such as principal components analysis, are used to determine ancestry in complex populations. Unfortunately, these data are not always available in the clinical and research settings. For example, electronic health records (EHRs), which are a rich source of temporal human disease data that could be used to enhance genetic studies, do not directly capture ancestry. Here, we present a novel algorithm for predicting genetic ancestry using only variables that are routinely captured in EHRs, such as self-reported race and ethnicity, and condition billing codes. Using patients that have both genetic and clinical information at Columbia University/ New York-Presbyterian Irving Medical Center, we developed a pipeline that uses only clinical data to predict the genetic ancestry of all patients of which more than 80% identify as other or unknown. Our ancestry estimates can be used in observational studies of disease inheritance, to guide genetic cohort studies, or to explore health disparities in clinical care and outcomes.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.