Back

Refined selection of individuals for preventive cardiovascular disease treatment with a Transformer-based risk model

Rao, S.; Li, Y.; Mamouei, M.; Salimi-Khorshidi, G.; Wamil, M.; Nazarzadeh, M.; Yau, C.; Collins, G.; Jackson, R.; Vickers, A.; Danaei, G.; Rahimi, K.

2024-09-26 cardiovascular medicine
10.1101/2024.09.25.24314371 medRxiv
Show abstract

ObjectiveTo develop and validate the Transformer-based Risk assessment survival model (TRisk), a novel deep learning model, for prediction of 10-year risk of cardiovascular disease (CVD) in both the general population and individuals with diabetes. DesignProspective open cohort study design. SettingPrimary and secondary care in England as provided by Clinical Practice Research Datalink (CPRD) Gold ParticipantsAn open cohort of 3 million adults aged 25 to 84 years was identified using linked primary and secondary electronic health records from 291 and 98 general practices in England and were used for model development and validation, respectively (i.e., general population cohort). Additionally, a second cohort of patients with diabetes was extracted. At study entry, patients in both cohorts were free of CVD and not prescribed statins. MethodsTRisk utilised all diagnosis, medication, procedure, and clinical test data up to study entry in linked longitudinal primary and secondary care electronic health records for prediction of 10-year risk of CVD. Discrimination, calibration, and decision curve analyses were conducted to investigate predictive performance. The proposed model was also compared against QRISK3 and a deep learning derivation model of QRISK3 (DeepSurv). Additional analyses compared discriminatory performance in other age groups, by sex, and across categories of socioeconomic status. Main outcome measuresIncident cardiovascular disease recorded in either linked general practice or hospital admission datasets provided by CPRD Gold. ResultsTRisk demonstrated superior discrimination (C-index in the general population: 0.910; 95% confidence interval [CI]: 0.906 to 0.913). TRisks performance was found to be less sensitive to population age range than the benchmark models and outperformed other models also in analyses stratified by age, sex or socioeconomic status. All models were overall well-calibrated. In decision curve analyses, TRisk demonstrated greater net benefit than benchmark models across the range of relevant thresholds. At both the recommended 10% risk threshold and the 15% risk threshold, TRisk reduced both the total number of patients classified at high risk (by 22% and 35% respectively) and the number of false negatives as compared with currently recommended strategies. TRisk similarly outperformed other models in patients with diabetes. Compared with the widely recommended treat-all policy approach for patients with diabetes, TRisk at a 10% risk threshold would lead to deselection of 24% of individuals with a small fraction of false negatives (0.2% of cohort). ConclusionTRisk enabled a more targeted selection of individuals at risk of CVD compared to benchmark statistical and deep learning models, in both the general population and patients with diabetes. Incorporation of TRisk into routine clinical care would allow a reduction in the number of treatment-eligible patients by approximately one-third while preventing at least as many events as with currently adopted approaches.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.