Back

Enhanced Diabetes Prediction Using Novel Additive-Multiplicative Neural Networks: A Comprehensive Machine Learning Analysis of the PIMA Indians Dataset

Demirel, S.; Aytekin, K.; agraz, m.

2025-09-22 endocrinology
10.1101/2025.09.20.25336250 medRxiv
Show abstract

BackgroundEarly diabetes detection remains challenging, requiring robust machine learning approaches that balance accuracy with clinical interpretability for effective diagnostic support. MethodsWe are proposing a novel Additive and Multiplicative Neurons Network (AMNN) that combines both additive and multiplicative computational pathways to capture complex nonlinear relationships in diabetes prediction. Using the PIMA Indians Diabetes dataset (n=768), we compared AMNN against nine established algorithms including XGBoost, KAN, and traditional neural networks. Data preprocessing included SMOTE oversampling for class imbalance, and model interpretability was enhanced through SHAP and LIME explainable AI techniques. ResultsThe AMNN model outperformed all baseline approaches, achieving 75.76% accuracy, a 76.18% F1-score, and an AUC-ROC of 0.8206. Across both traditional feature selection techniques and explainable AI analyses, glucose levels, BMI, age, and pregnancy count consistently emerged as the most influential predictors. ConclusionsThe AMNN framework demonstrates strong potential for diabetes prediction by balancing accuracy with clinical interpretability. The key predictors it highlights align closely with established medical knowledge, reinforcing confidence in its outputs and suitability for use in clinical decision-making workflows. This hybrid neural network approach represents a promising step toward transparent, AI-assisted diagnostic tools that can support healthcare professionals in practice.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.