Enhancing Clinical Utility of AI Models in PCOS: Integrating Staged Modelling with uncertainty aware Triage and Conformal Prediction for Cost-Efficient Risk Stratification
Adegoke, A. A.; Babalola, I.; Odesola, P. A.
Show abstract
Polycystic Ovary Syndrome (PCOS) is a common endocrine disorder affecting women of reproductive age, yet its diagnosis remains challenging due to reliance on costly investigations and limited diagnostic capacity in many clinical settings. We develop and assess a two-stage modelling pipeline in which Stage 1 uses low-cost demographic and routine clinical variables for initial screening, and Stage 2 augments these with laboratory and ultrasound features when escalation is warranted. Logistic Regression (LR) and Random Forest (RF) models are evaluated using calibration and classification metrics (AUC, accuracy, F1, precision, recall, Brier score, ECE), alongside clinical utility via decision curve analysis (DCA). To enhance safety and transparency, we integrate conformal prediction (CP) to provide finite-sample coverage guarantees and controlled abstention. Across train-test performance and out-of-fold evaluations, both models demonstrated consistent performance gains from Stage 1 to Stage 2. AUC increased by 6.9% for LR and increased by 7.4% for RF. LR exhibited more favourable calibration in most settings, while RF achieved higher precision, particularly after escalation. DCA showed higher net benefit for Stage 2 across clinically relevant thresholds. Feature sensitivity analysis indicated that a compact subset of inexpensive predictors preserved over 80% of maximal AUC, supporting cost-efficient screening. CP achieved 94.5% overall coverage with a 41.3% abstention rate, maintaining near-nominal validity across age and BMI subgroups. Capacity-constrained triage experiments showed that prioritising highest-risk cases maximised net benefit when resources were limited. These findings suggest that staged, uncertainty-informed modelling may offer practical steps toward clinically aligned and resource-aware AI support for PCOS assessment.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 92%
- A human-in-the-loop explanation framework for morphologically transparent AI predictions from whole-slide images 91%
- Assessment of Menstrual Health Status and Evolution through Mobile Apps for Fertility Awareness 91%
Similar papers in this journal
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 92%
- Assessing generalizability of an AI-based visual test for cervical cancer screening 92%
- Racial disparities in continuous glucose monitoring-based 60-min glucose predictions among people with type 1 diabetes 92%
Similar papers in this journal
- Subtyping of common complex diseases and disorders by integrating heterogeneous data. Identifying clusters among women with lower urinary tract symptoms in the LURN study 94%
- Predicting preterm births from electrohysterogram recordings via deep learning 92%
- ChatGPT-Enhanced ROC Analysis (CERA): A Shiny Web Tool for Finding Optimal Cutoff in Biomarker Analysis 92%
Similar papers in this journal
- High-throughput Phenotyping with Temporal Sequences 91%
- Causal modeling of chronic kidney disease in a participatory framework for informing the inclusion of social drivers in health algorithms 91%
- Learning from local to global - an efficient distributed algorithm for modeling time-to-event data 91%
Similar papers in this journal
- Mitigating Machine Learning Bias Between High Income and Low-Middle Income Countries for Enhanced Model Fairness and Generalizability 93%
- Machine learning for classifying chronic kidney disease and predicting creatinine levels using at-home measurements 93%
- Widely accessible prognostication using medical history for fetal growth restriction and small for gestational age in nationwide insured women 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.