Back

Development and Validation of a Multivariable Risk Prediction Model for Sudden Cardiac Death after Myocardial Infarction (PROFID Risk Model): Study Rationale, Design and Protocol

Martin, G. P.; Hindricks, G.; Akbarov, A.; Kapacee, Z.; Parkes, L. M.; Motmedi-Ghahfarokhi, G.; Ng, S.; Sprague, D.; Taleb, Y.; Ong, M.; Longato, E.; Miller, C. A.; Shamloo, A. S.; Albert, C.; Barthel, P.; Boveda, S.; Braunschweig, F.; Johansen, J. B.; Cook, N.; de Chillou, C.; Elders, P.; Faxen, J.; Friede, T.; Fusini, L.; Gale, C. P.; Jarkovsky, J.; Jouven, X.; Junttila, J.; Kiviniemi, A.; Kutyifa, V.; Lee, D.; Leigh, J.; Lenarczyk, R.; Leyva, F.; Maeng, M.; Manca, A.; Marijon, E.; Marschall, U.; Vinayagamoorthy, M.; Nielsen, J. C.; Olsen, T.; Pester, J.; Pontone, G.; Schmidt, G.; Schwartz,

2021-07-16 cardiovascular medicine
10.1101/2021.07.12.21260002 medRxiv
Show abstract

IntroductionSudden cardiac death (SCD) is the leading cause of death in patients with myocardial infarction (MI) and can be prevented by the implantable cardioverter defibrillator (ICD). Currently, risk stratification for SCD and decision on ICD implantation are based solely on impaired left ventricular ejection fraction (LVEF). However, this strategy leads to over- and under-treatment of patients because LVEF alone is insufficient for accurate assessment of prognosis. Thus, there is a need for better risk stratification. This is the study protocol for developing and validating a prediction model for risk of SCD in patients with prior MI. Methods and AnalysisThe EU funded PROFID project will analyse 23 datasets from Europe, Israel and the US ([~]225,000 observations). The datasets include patients with prior MI or ischemic cardiomyopathy with reduced LVEF<50%, with and without a primary prevention ICD. Our primary outcome is SCD in patients without an ICD, or appropriate ICD therapy in patients carrying an ICD as a SCD surrogate. For analysis, we will stack 18 of the datasets into a single database (datastack), with the remaining analysed remotely for data governance reasons (remote data). We will apply 5 analytical approaches to develop the risk prediction model in the datastack and the remote datasets, all under a competing risk framework: 1) Weibull model, 2) flexible parametric survival model, 3) random forest, 4) likelihood boosting machine, and 5) neural network. These dataset-specific models will be combined into a single model (one per analysis method) using model aggregation methods, which will be externally validated using systematic leave-one-dataset-out cross-validation. Predictive performance will be pooled using random effects meta-analysis to select the model with best performance. Ethics and disseminationLocal ethical approval was obtained. The final model will be disseminated through scientific publications and a web-calculator. Statistical code will be published through open-source repositories.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.