Risk stratification for the rapid pain progression phenotype in knee osteoarthritis using interpretable multimodal machine learning: Development in the Osteoarthritis Initiative and external evaluation in the Prospective Cohort of Osteoarthritis from A Coruna
J Blanco, F. J.; Martinez-Sotodosos, L.; Oreiro, N.; Galindo, L.; Vazquez-Garcia, J.; Noriega-Cobo, D. M.; Lourido, L.; Paz-Gonzalez, R.; Quaranta, P.; Calamia, V.; Ruiz-Romero, C.; Rego-Perez, I.
Show abstract
Objective. To develop an interpretable multimodal machine-learning model for risk stratification of the rapid pain progression phenotype in knee osteoarthritis and to evaluate its performance in the independent PROCOAC cohort. Methods. An elastic-net logistic regression model was trained using Osteoarthritis Initiative (OAI) data. Rapid pain progression was defined over overlapping 24-month windows using normalized WOMAC pain. Harmonized clinical, genetic and proteomic candidates were evaluated, with feature selection by permutation importance. The frozen algorithm was tested in an OAI hold-out set and externally evaluated in PROCOAC. Logistic recalibration corrected prevalence shifts. Clinical utility was assessed by decision curve analysis. Results. OAI comprised 2,934 individuals and 14,488 instances. Feature pruning reduced 159 candidates to a 19-variable clinical-genetic signature driven by Kellgren-Lawrence grade, localized knee pain, BMI and two genetic variants (rs73631790, rs9912678); no proteomic variable was retained. External testing in PROCOAC (582 individuals, 1609 instances) showed ROC-AUC 0.744 (95% CI 0.714 to 0.772) and PR-AUC 0.519. Following recalibration, the sensitive screening threshold yielded NPV 0.875 (95% CI 0.849 to 0.898) and sensitivity 0.804 (95% CI 0.760 to 0.844), whereas the high-specificity threshold achieved PPV 0.610 (95% CI 0.523 to 0.692) and specificity 0.941 (95% CI 0.924 to 0.954). Decision curve analysis showed positive net benefit at both thresholds, supporting a three-tier risk stratification framework. Conclusions. This externally evaluated model identified patients at risk of rapid pain progression using an MRI-free clinical-genetic signature. Recalibrated thresholds may support risk-adapted monitoring, advanced imaging prioritization and trial enrichment.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Digital health technologies and machine learning augment patient reported outcomes to remotely characterise rheumatoid arthritis 93%
- Cross-Platform Omics Prediction procedure enables precision medicine in patients with stage-III melanoma 92%
- Federated Target Trial Emulation using Distributed Observational Data for Treatment Effect Estimation 92%
Similar papers in this journal
- Multi-ancestry omic Mendelian randomization revealing putative drug targets of COVID-19 severity 92%
- Integrative deep learning analysis improves colon adenocarcinoma patient stratification at risk for mortality 91%
- Machine learning guided association of adverse drug reactions with in vitro target-based pharmacology 91%
Similar papers in this journal
- Antibody therapy reverses biological signatures of COVID-19 progression 92%
- Determination of permissive and restraining cancer-associated fibroblast (DeCAF) subtypes 92%
- Enhancer profiling identifies epigenetic markers of endocrine resistance and reveals therapeutic options for metastatic castration-resistant prostate cancer patients 91%
Similar papers in this journal
- Single-cell chromatin and transcriptome dynamics of Synovial Fibroblasts transitioning from homeostasis to pathology in modelled TNF-driven arthritis 92%
- scDrugPrio: A framework for the analysis of single-cell transcriptomics to address multiple problems in precision medicine in immune-mediated inflammatory diseases 92%
- Refining epigenetic prediction of chronological and biological age 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.