Missingness Adapted Group Informed Clustered (MAGIC)-LASSO: A novel paradigm for prediction in data with widespread non-random missingness
Gentry, A. E.; Kirkpatrick, R. M.; Peterson, R. E.; Webb, B. T.
Show abstract
The availability of large-scale biobanks linking rich phenotypes and biological measures is a powerful opportunity for scientific discovery. However, real-world collections frequently have extensive non-random missingness. While missing data prediction is possible, performance is significantly impaired by block-wise missingness inherent to many biobanks. To address this, we developed Missingness Adapted Group-wise Informed Clustered (MAGIC)-LASSO which performs hierarchical clustering of variables based on missingness followed by sequential Group LASSO within clusters. Variables are pre-filtered for missingness and balance between training and target sets with final models built using stepwise inclusion of features ranked by completeness. This research has been conducted using the UK Biobank (n>500k) to predict unmeasured Alcohol Use Disorders Identification Test (AUDIT) scores. The phenotypic correlation between measured and predicted total score was 0.67 while genetic correlations between independent subjects was high >0.86, demonstrating the method has significant accuracy and utility.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Comprehensive Evaluation of Methods for Mendelian Randomization Using Realistic Simulations and an Analysis of 38 Biomarkers for Risk of Type-2 Diabetes 93%
- Bias in two-sample Mendelian randomization when using heritable covariable-adjusted summary associations 92%
- Causes of Outcome Learning: A causal inference-inspired machine learning approach to disentangling common combinations of potential causes of a health outcome 91%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Analyzing Biomarker Discovery: Estimating the Reproducibility of Biomarker Sets 93%
- Towards development of a statistical framework to evaluate myotonic dystrophy type 1 mRNA biomarkers in the context of a clinical trial 92%
- Assessing the performance of genome-wide association studies for predicting disease risk 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.