Explainable AI to predict a complex multifactorial outcome, childhood obesity: Application to clinical epidemiology
Chen, F.; Melton, P.; Vinsen, K.; Mori, T. A.; Beilin, L.; Huang, R.-C.
Show abstract
BackgroundChildhood obesity, driven by genetic and epidemiological factors, poses significant health risks, yet traditional machine learning models lack interpretability for clinical use. ObjectiveThis study aims to apply Kolmogorov-Arnold Networks (KAN), an explainable machine learning model, to predict body mass index (BMI) at age 8 as an indicator of obesity risk and to develop a publicly accessible prediction tool. MethodsWe utilized the Raine Study Gen2 cohort (n=2,868) to train KAN and traditional models (such as Random Forest, Gradient Boosting, Lasso, and Multi-Layer Perceptron) using perinatal, early-life, and polygenic risk score (PGS) data collected before age 5. Feature importance was analyzed across all the models. A publicly accessible online calculator was developed for practical use. ResultsKAN achieved an R2 of 0.81, outperforming traditional models. Key predictors included Year 5 BMI z-score, mid-arm circumference, occupation of mother, and PGS. The online calculator supports predictions without PGS, maintaining an R2 of 0.81. ConclusionsKANs transparent formulas enhance interpretability, offering a practical approach to predicting childhood obesity. The freely accessible tool enables clinicians to implement personalized prevention strategies, advancing precision medicine. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=186 HEIGHT=200 SRC="FIGDIR/small/25330041v3_ufig1.gif" ALT="Figure 1"> View larger version (59K): org.highwire.dtl.DTLVardef@1551d55org.highwire.dtl.DTLVardef@f8e337org.highwire.dtl.DTLVardef@d5447org.highwire.dtl.DTLVardef@11841f8_HPS_FORMAT_FIGEXP M_FIG KAN model predicts childhood obesity (BMI at age 8), showcasing key features, top performance, and accurate formularised results with epidemiological and genetic factors. Online calculator is available at https://bmi-y8-calc.onrender.com/. C_FIG
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predicting long-term Type 2 Diabetes with Support Vector Machine using Oral Glucose Tolerance Test 94%
- ChatGPT-Enhanced ROC Analysis (CERA): A Shiny Web Tool for Finding Optimal Cutoff in Biomarker Analysis 93%
- Optimization of nutritional strategies using a mechanistic computational model in prediabetes: Application to the J-DOIT1 study data 93%
Similar papers in this journal
- Racial disparities in continuous glucose monitoring-based 60-min glucose predictions among people with type 1 diabetes 93%
- Machine learning-based equations for improved body composition estimation in Indian adults 93%
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 93%
Similar papers in this journal
- Bayesian Structural Time Series for Biomedical Sensor Data: A Flexible Modeling Framework for Evaluating Interventions 93%
- Deep learning approach for automatic assessment of schizophrenia and bipolar disorder in patients using R-R intervals 93%
- Model guided trait-specific co-expression network estimation as a new perspective for identifying molecular interactions and pathways 92%
Similar papers in this journal
- Machine learning for classifying chronic kidney disease and predicting creatinine levels using at-home measurements 94%
- Mitigating Machine Learning Bias Between High Income and Low-Middle Income Countries for Enhanced Model Fairness and Generalizability 93%
- Stochastic LASSO for extremely high-dimensional genomic data 93%
Similar papers in this journal
- A dynamic ensemble model for short-term forecasting in pandemic situations 91%
- Measuring the impact of nonpharmaceutical interventions on the SARS-CoV-2 pandemic at a city level: An agent-based computational modeling study of the City of Natal 90%
- Influence of parental anthropometry and gestational weight gain on intrauterine growth and neonatal outcomes: Findings from the MAI cohort study in rural India 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.