Classifying polyneuropathy and myopathy patients on Electronic Health Records
Ahmed, M. S.; Truong, N. D. K.; Nyoungui, E.; Zhao, J.; Wedemeyer, H.; Mayer, R.; Schuster, V.; Zschüntzsch, J.; Röttger, R.
Show abstract
BackgroundRare neuromuscular diseases such as polyneuropathy (PN) and myopathy (MY) often share symptomatic characteristics, leading to diagnostic challenges and delays. Machine learning applied to routine care data of electronic health records (EHRs) offers the potential for accelerating accurate diagnosis. ObjectiveTo develop and evaluate machine learning models to distinguish between patients with PN and MY using EHR data, as a step toward tools that could support improved diagnostic processes. MethodsWe analyzed EHR data from 2,181 patients (1,853 PN, 328 MY) provided by the Medical Data Integration Center of the University of Gottingen. The features were curated according to the recommendations of the physicians, the literature, and statistical analysis. We implemented Logistic Regression, Random Forest, and XGBoost models, optimized with Grid Search, and addressed class imbalance using SMOTE. ResultsRandom Forest and XGBoost models achieved the best performance with F1 Macro scores of 0.82-0.84 and AUC-ROC scores of 0.92-0.93 when trained on demographic data, feature-engineered variables, laboratory test results, and ICD-10 codes. Patient age emerged as a significant predictive factor, with MY patients typically diagnosed at younger ages (mean=51.39) than PN patients (mean=67.11). ConclusionMachine learning models can effectively differentiate between PN and MY patients using EHR data with low data depth, potentially accelerating diagnostic processes for these rare neuromuscular diseases. Availability and implementationhttps://gitlab.sdu.dk/screen4care/classifying-pn-and-my
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Characterizing subgroup performance of probabilistic phenotype algorithms within older adults: A case study for dementia, mild cognitive impairment, and Alzheimer’s and Parkinson’s diseases 94%
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 94%
- Modeling physician variability to prioritize relevant medical record information 93%
Similar papers in this journal
- AI-MET: A Deep Learning-based Clinical Decision Support System for Distinguishing Multisystem Inflammatory Syndrome in Children from Endemic Typhus 95%
- Unsupervised Discovery of Risk Profiles on Negative and Positive COVID-19 Hospitalized Patients 94%
- SymScore: Machine Learning Accuracy Meets Transparency in a Symbolic Regression-Based Clinical Score Generator 94%
Similar papers in this journal
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 95%
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 93%
- Racial disparities in continuous glucose monitoring-based 60-min glucose predictions among people with type 1 diabetes 93%
Similar papers in this journal
- Prediction of Sepsis Mortality in ICU Patients Using Machine Learning Methods 94%
- OASIS+: leveraging machine learning to improve the prognostic accuracy of OASIS severity score for predicting in-hospital mortality 94%
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 94%
Similar papers in this journal
- High-throughput Phenotyping with Temporal Sequences 94%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 93%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.