Back

Algorithms for the identification of prevalent diabetes in the All of Us Research Program validated using polygenic scores - a new resource for diabetes precision medicine

Szczerbinski, L.; Mandla, R.; Schroeder, P.; Porneala, B. C.; Li, J. H.; Florez, J. C.; Mercader, J. M.; Manning, A. K.; Udler, M. S.

2023-09-05 endocrinology
10.1101/2023.09.05.23295061 medRxiv
Show abstract

OBJECTIVEThe study aimed to develop and validate algorithms for identifying people with type 1 and type 2 diabetes in the All of Us Research Program (AoU) cohort, using electronic health record (EHR) and survey data. RESEARCH DESIGN AND METHODSTwo sets of algorithms were developed, one using only EHR data (EHR), and the other using a combination of EHR and survey data (EHR+). Their performance was evaluated by testing their association with polygenic scores for both type 1 and type 2 diabetes. RESULTSFor type 1 diabetes, the EHR-only algorithm showed a stronger association with T1D polygenic score (p=3x10-5) than the EHR+. For type 2 diabetes, the EHR+ algorithm outperformed both the EHR-only and the existing AoU definition, identifying additional cases (25.79% and 22.57% more, respectively) and showing stronger association with T2D polygenic score (DeLong p=0.03 and 1x10-4, respectively). CONCLUSIONSWe provide new validated definitions of type 1 and type 2 diabetes in AoU, and make them available for researchers. These algorithms, by ensuring consistent diabetes definitions, pave the way for high-quality diabetes research and future clinical discoveries. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=62 SRC="FIGDIR/small/23295061v1_ufig1.gif" ALT="Figure 1"> View larger version (20K): org.highwire.dtl.DTLVardef@192a4e2org.highwire.dtl.DTLVardef@86fe8borg.highwire.dtl.DTLVardef@b18e50org.highwire.dtl.DTLVardef@f6496b_HPS_FORMAT_FIGEXP M_FIG C_FIG Article Highlightsa. Why did we undertake this study?This study was conducted to develop and validate algorithms for identifying type 1 and type 2 diabetes cases in the All of Us Research Program (AoU). b. What is the specific question(s) we wanted to answer?Can accurate algorithms for type 1 and type 2 diabetes identification be developed and validated using AoU cohort Electronic Health Record (EHR) and survey data? Do the identified diabetes cases show association with polygenic scores in diverse populations? c. What did we find?We developed a new validated type 1 diabetes definition and expanded upon the existing type 2 diabetes definition. d. What are the implications of our findings?The developed algorithms can be universally implemented in AoU for identifying study participants for well-defined case-control diabetes studies.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.