Identifying type 1 and 2 diabetes in population level data: assessing the accuracy of published approaches
Thomas, N. J. M.; McGovern, A.; Young, K. G.; Sharp, S.; Weedon, M.; Hattersley, A.; Dennis, J.; Jones, A.
Show abstract
AimsPopulation datasets are increasingly used to study type 1 or 2 diabetes, and inform clinical practice. However, correctly classifying diabetes type, when insulin treated, in population datasets is challenging. Many different approaches have been proposed, ranging from simple age or BMI cut offs, to complex algorithms, and the optimal approach is unclear. We aimed to compare the performance of approaches for classifying insulin treated diabetes for research studies, evaluated against two independent biological definitions of diabetes type. MethodWe compared accuracy of thirteen reported approaches for classifying insulin treated diabetes into type 1 and type 2 diabetes in two population cohorts with diabetes: UK Biobank (UKBB) n=26,399 and DARE n=1,296. Overall accuracy and predictive values for classifying type 1 and 2 diabetes were assessed using: 1) a type 1 diabetes genetic risk score and genetic stratification method (UKBB); 2) C-peptide measured at >3 years diabetes duration (DARE). ResultsAccuracy of approaches ranged from 71%-88% in UKBB and 68%-88% in DARE. All approaches were improved by combining with requirement for early insulin treatment (<1 year from diagnosis). When classifying all participants, combining early insulin requirement with a type 1 diabetes probability model incorporating continuous clinical features (diagnosis age and BMI only) consistently achieved high accuracy, (UKBB 87%, DARE 85%). Self-reported diabetes type alone had high accuracy (UKBB 87%, DARE 88%) but was available in just 15% of UKBB participants. For identifying type 1 diabetes with minimal misclassification, using models with high thresholds or young age at diagnosis (<20 years) had the highest performance. An online tool developed from all UKBB findings allows the optimum approach of those tested to be selected based on variable availability and the research aim. ConclusionSelf-reported diagnosis and models combining continuous features with early insulin requirement are the most accurate methods of classifying insulin treated diabetes in research datasets without measured classification biomarkers.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Relation of incident Type 1 diabetes to recent COVID-19 infection: cohort study using e-health record linkage in Scotland 96%
- Accounting for age-related increases in HbA1c more accurately quantifies risk of Type 1 Diabetes progression in islet autoantibody-positive adults 94%
- Early metabolic features of genetic liability to type 2 diabetes: cohort study with repeated metabolomics across early life 93%
Similar papers in this journal
- Treatment outcomes with oral anti-hyperglycaemic therapies in people with diabetes secondary to a pancreatic condition (type 3c diabetes): A population-based cohort study 96%
- Covid-19 fatality prediction in people with diabetes and prediabetes using a simple score at hospital admission 93%
- Subgroups of adult-onset diabetes: a prospective follow-up study of progression of insulin resistance and deficiency and association with liver steatosis and fibrosis 92%
Similar papers in this journal
- Ethnic differences in the manifestation of early-onset type 2 diabetes 96%
- CREBRF missense variant rs373863828 has both direct and indirect effects on type 2 diabetes and fasting glucose in Polynesians living in Samoa and Aotearoa New Zealand 94%
- Comprehensive validation of fasting- and oral glucose tolerance test-based indices of insulin secretion against gold-standard measures 93%
Similar papers in this journal
- All thresholds of maternal hyperglycaemia from the WHO 2013 criteria for gestational diabetes identify women with a higher genetic risk for type 2 diabetes 93%
- What’s UPDOG? A novel tool for trans-ancestral polygenic score prediction 87%
- The consequences of adjustment, correction and selection in genome-wide association studies used for two-sample Mendelian randomization 86%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.