Machine-learning approach for Type 2 Diabetes diagnosis and prognosis models over heterogeneous feature spaces
Navarro-Cerdan, J. R.; Pons-Suner, P.; Arnal, L.; Arlandis, J.; Llobet, R.; Perez-Cortes, J.-C.; Lara-Hernandez, F.; Alvarez, L.; Garcia-Garcia, A.-B.; Chaves, F. J.
Show abstract
This research aims to evaluate the Type 2 Diabetes (T2D) diagnosis and prognosis power from heterogeneous environmental, lifestyle and biochemistry data. Model estimation has previously addressed three main actions as: 1) Missingvalue imputation using specific univariant and multivariant imputers accommodated to each particular feature; 2) Quasi-constancy detection in variables; 3) Constructing geographical pollution and rent data from municipality information. Next, different T2D diagnosis and prognosis models are fitted and evaluated, showing increasing performance as more specific features become available while the prediction cost rises as a consequence of requiring more specific data. Finally, four models are obtained: two of them for T2D diagnosis and the other two for T2D prognosis respectively, with performances ranging from 73.3 to 95.41 AUC-ROC. One pair of diagnosis and prognosis models were thought for a global testing that can be done in general locations by only asking general lifestyle-related questions. On the other hand, the other pair, which achieves higher performances, is thought to be applied in a clinical environment where it is easy to obtain more specific biochemistry measures.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Finding disease modules for cancer and COVID-19 in gene co-expression networks with the Core&Peel method 94%
- Selecting the most important self-assessed features for predicting conversion to Mild Cognitive Impairment with Random Forest and Permutation-based methods 93%
- Testing the phenotypic decanalization hypothesis: social determinants of hyperglycemia and type 2 diabetes in adult urban Argentinian population 93%
Similar papers in this journal
Similar papers in this journal
- A Utility-Based Machine Learning-Driven Personalized Lifestyle Recommendation for Cardiovascular Disease Prevention 93%
- Computational Strategies in Nutrigenetics: Constructing a Reference Dataset of Nutrition-Associated Genetic Polymorphisms 91%
- A methodology of phenotyping ICU patients from EHR data: high-fidelity, personalized, and interpretable phenotypes estimation 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.