Back

Improving Diagnostic Sensitivity for Imbalanced Musculoskeletal Disorder Data: A Sensitivity-Based Multi-Sampling Technique for Osteoarthritis Prediction

Kim, J.-h.

2023-11-20 rehabilitation medicine and physical therapy
10.1101/2023.11.19.23298738 medRxiv
Show abstract

BackgroundMedical datasets containing musculoskeletal disorders may have data imbalances due to the incidence of the disease, which may limit the predictive ability, such as the sensitivity, of musculoskeletal diagnostic prediction models built from these data. This study aimed to increase the sensitivity performance of osteoarthritis (OA) prediction when building a model by adjusting an OA imbalanced dataset using a sensitivity-based multi-sampling (SMS) technique. MethodsOA Data were obtained from the Korea National Health and Nutrition Examination Survey (KNHANES). SMS technique combining oversampling and undersampling was applied to the imbalanced OA data, and the RandomForest algorithm was used for machine learning modeling. Model performance was evaluated based on accuracy, sensitivity, and specificity and compared with other hybrid sampling techniques. ResultIn the SMS technique, ADASYN, Borderline-SMOTE, SMOTE oversampling and ENN undersampling techniques were combined and applied. The OA prediction model using the SMS technique showed the highest sensitivity (82.20) but the lowest specificity (82.26) and accuracy (82.26) compared to other hybrid models. ConclusionSMS technology offers a potential solution for improving sensitivity performance for prediction models built on medical data imbalances due to low-incidence diseases. Nonetheless, caution is warranted due to the concern that while improving sensitivity, it may decrease specificity with a trade-off.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.