Discriminating the prodromal stage of multiple sclerosis using longitudinal health administrative claims data and machine learning-based sequence analysis
Klempir, O.; Hola, M.; Rozanek, M.; Mullerova, J. G.; Tichopad, A.
Show abstract
BackgroundMultiple sclerosis (MS) is a chronic autoimmune disease of the central nervous system. Early detection of the prodromal phase could enable timely interventions to potentially modify disease progression. This study leverages longitudinal health administrative claim (HAC) data to identify patterns distinguishing the prodromal stage of MS from other neurological conditions. MethodsHAC data from the Czech Health Insurance Bureau (2017-2022) was analyzed across three cohorts: a target MS cohort with confirmed diagnoses, a control cohort with inconsistent MS suspicions, and a cohort with related disorders. For healthcare utilization and diagnostic code data representation, we employed two approaches: temporal analysis using various time windows relative to the index date (including pre- and post-index date comparisons) and a separate segment-based analysis. Features were extracted using token frequencies and word embeddings. Random forest models were evaluated using Area Under the Receiver Operating Characteristic Curve (AUC) to assess performance. ResultsEach cohort included several hundred to over a thousand individuals. The models achieved AUCs around 0.9 for distinguishing the target cohort from controls, with even higher performance in differentiating pre- and post-diagnosis phases. Longer observation windows enhanced predictive accuracy, and feature extraction methods like TF-IDF and word2vec yielded the most consistent results. Segment-based analysis identified a subset of individuals for potential diagnostic reclassification. Interpretable machine learning techniques were integrated into the analysis pipeline. ConclusionsThis study highlights the potential of HAC data for detecting early prodromal indicators of MS. Unlike previous research, which often focused on the volume of healthcare utilization, this work explores the informational content within diagnostic codes and healthcare utilization patterns. The findings align with existing research on early neurological condition detection, demonstrating that administrative data could support early identification and intervention in MS and possibly other diseases.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Conformal prediction enables disease course prediction and allows individualized diagnostic uncertainty in multiple sclerosis 94%
- COVID-19 diagnosis prediction by symptoms of tested individuals: a machine learning approach 93%
- Quantifying Device Type and Handedness Biases in a Remote Parkinson’s Disease AI-Powered Assessment 92%
Similar papers in this journal
- AI-based mining of biomedical literature: Applications for drug repurposing for the treatment of dementia 93%
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 93%
- Deep ensemble multitask classification of emergency medical call incidents combining multimodal data improves emergency medical dispatch 93%
Similar papers in this journal
Similar papers in this journal
- Machine-learning-based prediction of disability progression in multiple sclerosis: an observational, international, multi-center study 97%
- Explainable deep learning for disease activity prediction in chronic inflammatory joint diseases 93%
- Identification of predictive patient characteristics for assessing the probability of COVID-19 in-hospital mortality 93%
Similar papers in this journal
- Brain predictors of fatigue in Rheumatoid Arthritis: a machine learning study 93%
- Imbalanced Machine Learning Classification Models For Removal Biosimilar Drugs And Increased Activity In Patients With Rheumatic Diseases 93%
- Towards development of a statistical framework to evaluate myotonic dystrophy type 1 mRNA biomarkers in the context of a clinical trial 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.