Back

Classifying polyneuropathy and myopathy patients on Electronic Health Records

Ahmed, M. S.; Truong, N. D. K.; Nyoungui, E.; Zhao, J.; Wedemeyer, H.; Mayer, R.; Schuster, V.; Zschüntzsch, J.; Röttger, R.

2025-12-12 health informatics
10.64898/2025.12.11.25342051 medRxiv
Show abstract

BackgroundRare neuromuscular diseases such as polyneuropathy (PN) and myopathy (MY) often share symptomatic characteristics, leading to diagnostic challenges and delays. Machine learning applied to routine care data of electronic health records (EHRs) offers the potential for accelerating accurate diagnosis. ObjectiveTo develop and evaluate machine learning models to distinguish between patients with PN and MY using EHR data, as a step toward tools that could support improved diagnostic processes. MethodsWe analyzed EHR data from 2,181 patients (1,853 PN, 328 MY) provided by the Medical Data Integration Center of the University of Gottingen. The features were curated according to the recommendations of the physicians, the literature, and statistical analysis. We implemented Logistic Regression, Random Forest, and XGBoost models, optimized with Grid Search, and addressed class imbalance using SMOTE. ResultsRandom Forest and XGBoost models achieved the best performance with F1 Macro scores of 0.82-0.84 and AUC-ROC scores of 0.92-0.93 when trained on demographic data, feature-engineered variables, laboratory test results, and ICD-10 codes. Patient age emerged as a significant predictive factor, with MY patients typically diagnosed at younger ages (mean=51.39) than PN patients (mean=67.11). ConclusionMachine learning models can effectively differentiate between PN and MY patients using EHR data with low data depth, potentially accelerating diagnostic processes for these rare neuromuscular diseases. Availability and implementationhttps://gitlab.sdu.dk/screen4care/classifying-pn-and-my

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.