Predictive Performance Precision Analysis in Medicine: Identification of low-confidence predictions at patient and profile levels (MED3pa I)
Lefebvre, O.; Camirand Lemyre, F.; Ethier, J.-F.; Chikouche, l. H.; Amriou, L.; Poenaru, D. D.; Vallieres, M.
Show abstract
ObjectiveArtificial Intelligence models are increasingly used in healthcare, yet global performance metrics can mask variations in reliability across individual patients or subgroups with shared attributes, called patient profiles. This study introduces MED3pa, a method that identifies when models are less reliable, allowing clinicians to better assess model limitations. Materials and MethodsWe propose a framework that estimates predictive confidence using three combined approaches: Individualized (IPC), Aggregated (APC), and Mixed Predictive Confidence (MPC). IPC estimates confidence for each patient, APC assesses it across profiles, and MPC combines both. We evaluate our method on four datasets: one simulated, two public, and one private clinical dataset. Metrics by Declaration Rate (MDR) curves show how performance changes when retaining only the most confident predictions, while interpretable decision trees reveal profiles with higher or lower model confidence. ResultsWe demonstrate our method in internal, temporal, and external validation settings, as well as through a clinical example. In internal validation, limiting predictions to the 93% most confident cases improved sensitivity by 14.3% and the AUC by 5.1%. In the clinical example, MED3pa identified a patient profile with high misclassification risk, demonstrating its potential for safer deployment. DiscussionBy identifying low-confidence predictions, our framework improves model reliability in clinical settings. It can be integrated into decision support systems to help clinicians make more informed decisions. Confidence thresholds help balance model performance with the proportion of patients for whom predictions are considered reliable. ConclusionBetter leveraging confidence in model predictions could improve reliability and trustworthiness, supporting safer and more effective use in healthcare.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 96%
- Developing Machine Learning Models for Predicting Intensive Care Unit Resource Use During the COVID-19 Pandemic 96%
- Emergency department admissions during COVID-19: explainable machine learning to characterise data drift and detect emergent health risks 96%
Similar papers in this journal
Similar papers in this journal
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 96%
- Confidence-based laboratory test reduction recommendation algorithm 95%
- OASIS+: leveraging machine learning to improve the prognostic accuracy of OASIS severity score for predicting in-hospital mortality 94%
Similar papers in this journal
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 96%
- Clinical Time-to-Event Prediction Enhanced by Incorporating Compatible Related Outcomes 94%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.