Predicting a diagnosis of ankylosing spondylitis using primary care health records: a machine learning approach.
Kennedy, J. I.; Kennedy, N.; Cooksey, R.; Choy, E.; Siebert, S.; Rahman, M.; Brophy, S. T.
Show abstract
Ankylosing spondylitis is the second most common cause of inflammatory arthritis. However, a successful diagnosis can take a decade to confirm from symptom onset (via x-rays). The aim of this study was to use machine learning methods to develop a profile of the characteristics of people who are likely to be given a diagnosis of AS in future. The Secure Anonymised Information Linkage databank was used. Patients with ankylosing spondylitis were identified using their routine data and matched with controls who had no record of a diagnosis of ankylosing spondylitis or axial spondyloarthritis. Data was analysed separately for men and women. The model was developed using feature/variable selection and principal component analysis to develop decision trees. The decision tree with the highest average F value was selected and validated with a test dataset. The model for men indicated that lower back pain, uveitis, and NSAID use under age 20 is associated with AS development. The model for women showed an older age of symptom presentation compared to men with back pain and multiple pain relief medications. The models showed good prediction (positive predictive value 70%-80%) in test data but in the general population where prevalence is very low (0.09% of the population in this dataset) the positive predictive value would be very low (0.33%-0.25%). Machine learning can be used to help profile and understand the characteristics of people who will develop AS, and in test datasets with artificially high prevalence, will perform well. However, when applied to a general population with low prevalence rates, such as that in primary care, the positive predictive value for even the best model would be 1.4%. Multiple models may be needed to narrow down the population over time to improve the predictive value and therefore reduce the time to diagnosis of ankylosing spondylitis.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Prodromal symptoms of rheumatoid arthritis in a primary care database: variation by ethnicity and socioeconomic status 94%
- Incorporating computer vision on smart phone photographs into screening for inflammatory arthritis: results from an Indian patient cohort 93%
- Risk of death among people with rare autoimmune diseases compared to the general population in England during the 2020 COVID-19 pandemic 93%
Similar papers in this journal
- UK osteopathic practice in 2019: a retrospective analysis of practice data 94%
- The economic burden of low back pain in KwaZulu-Natal, South Africa: a prevalence-based cost-of-illness analysis from the healthcare provider’s perspective 93%
- Imbalanced Machine Learning Classification Models For Removal Biosimilar Drugs And Increased Activity In Patients With Rheumatic Diseases 92%
Similar papers in this journal
Similar papers in this journal
- Versus Arthritis Musculoskeletal Disorders Research Advisory Group Priority Setting Exercise Protocol 93%
- Incidence and management of inflammatory arthritis in England before and during the COVID-19 pandemic: a population-level cohort study using OpenSAFELY 93%
- Gout and coronavirus disease-19 (COVID-19): the risk of diagnosis and death in the UK Biobank 91%
Similar papers in this journal
- Distributions of Recorded Pain in Mental Health Records: A Natural Language Processing Based Study 93%
- Prevalence and determinants of persistent symptoms after infection with SARS-CoV-2: Protocol for an observational cohort study (LongCOVID-study) 92%
- Tracking Persistent Symptoms in Scotland (TraPSS): A Longitudinal Prospective Cohort Study of COVID-19 Recovery After Mild Acute Infection 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.