Predicting the need for medical care after toxin exposure using SHAP-interpretable gradient boosting
Lerogeron, H.; Gueguen, L.; Chary, M.; Nguyen, K. A.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWO_ST_ABSObjectiveC_ST_ABSExperts in poison control centers must accurately and efficiently assess the severity of an exposure, neither delaying care nor pointlessly sending patients to the hospital, using only the information given during a first phone call. To help healthcare professionals (HP) make these difficult decisions, we developed and evaluated a machine learning-based algorithm that predicts whether a patient should seek medical help or not, based solely on the information provided during their first call to the poison control center, for all kinds of mono-intoxications. MethodsWe extracted data recorded by clinicians at the Lyon PCC between 2000 and 2025. Cases with missing original recommendations were excluded. We trained and compared several machine-learning models, emphasizing decision-tree-based and gradient-boosted tree approaches. Two classification tasks were defined: (1) binary triage (recommend emergency or non-emergency healthcare facility vs. stay at home) and (2) three-class triage (stay at home / non-emergency healthcare facility / emergency healthcare facility). Missing data were left as-is. Cross-validation and bootstrapping were used to ensure stable and statistically significant results. Model explainability was assessed with SHAP to identify the most important features for predictions. Model performance was evaluated using F1-score and ROC AUC; class imbalance was addressed during training. We compared our results to published algorithms that focus on single-substance intoxications. ResultsAfter processing, 220,825 cases remained. Recommended dispositions were: stay at home 66.6%, emergency facility 25.4%, and non-emergency facility 7.4%. For the binary task, XGBoost achieved the best performance (F1 = 0.748; ROC AUC = 0.820). For the three-class task, XGBoost again performed best (macro F1 = 0.657; multiclass ROC AUC = 0.859). The delay from exposure to call, SNOMED symptom codes, and the circumstance of exposure were the most influential features. Our results were competitive with algorithms focusing on intoxication due to a single substance. ConclusionGradient-boosted tree models can produce accurate, interpretable, and clinically relevant predictions of poisoning severity from routine PCC data. With external validation and prospective testing, such tools could complement expert judgment to improve triage consistency and patient outcomes.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Improving irregular temporal modeling by integrating synthetic data to the electronic medical record using conditional GANs: a case study of fluid overload prediction in the intensive care unit 94%
- Predicting the physiological effects of multiple drugs using electronic health record 93%
- Machine Learning Interpretability Methods to Characterize the Importance of Hematologic Biomarkers in Prognosticating Patients with Suspected Infection 93%
Similar papers in this journal
- Detecting Rare Diseases in Electronic Health Records Using Machine Learning and Knowledge Engineering: Case Study of Acute Hepatic Porphyria 93%
- Identification of high-risk COVID-19 patients using machine learning 92%
- ChatGPT-Enhanced ROC Analysis (CERA): A Shiny Web Tool for Finding Optimal Cutoff in Biomarker Analysis 92%
Similar papers in this journal
- Modeling physician variability to prioritize relevant medical record information 92%
- Characterizing subgroup performance of probabilistic phenotype algorithms within older adults: A case study for dementia, mild cognitive impairment, and Alzheimer’s and Parkinson’s diseases 92%
- Development and Evaluation of Machine Learning Models for the Detection of Emergency Department Patients with Opioid Misuse from Clinical Notes 91%
Similar papers in this journal
- Deep learning approach for automatic assessment of schizophrenia and bipolar disorder in patients using R-R intervals 92%
- Predicting the causative pathogen among children with pneumonia using a causal Bayesian network 92%
- An expert judgment model to predict early stages of the COVID-19 outbreak in the United States 92%
Similar papers in this journal
- Measuring the impact of nonpharmaceutical interventions on the SARS-CoV-2 pandemic at a city level: An agent-based computational modeling study of the City of Natal 91%
- Drug Supply Management at First-level Public Health Facilities: Case of Pyay District, Myanmar 90%
- Low-cost, local production of a safe and effective disinfectant for resource-constrained communities 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.