Back

Multitask Learning of Longitudinal Circulating Biomarkers and Clinical Outcomes: Identification of Optimal Machine-Learning and Deep-Learning Models

Yuan, M.; Su, S.; Ding, H.; Yang, Y.; Gupta, M.; Xu, X. S.

2023-08-21 pharmacology and toxicology
10.1101/2023.08.19.553991 bioRxiv
Show abstract

Many circulating biomarkers are assessed at different time intervals during clinical studies. Despite of the success of standard joint models in predicting clinical outcomes using low-dimensional longitudinal data (1-2 biomarkers), significant computational challenges are encountered when applying these techniques to high-dimensional biomarker datasets. Modern machine- or deep-learning models show potential for multiple biomarker processes, but systematic evaluations and applications to high-dimensional data in the clinical settings have yet to be reported. We aimed to enhance the scalability of joint modeling and provide guidance on optimal approaches for high-dimensional biomarker data and outcomes. We evaluated multiple deep-learning and machine-learning models using 24 clinical biomarkers and survival data from the SQUIRE trial, a phase 3 randomized clinical trial investigating necitumumab and standard gemcitabine/cisplatin treatment in patients with squamous non-small-cell lung cancer (NSCLC). Overall, we confirmed that longitudinal models enabled more accurate prediction of patients survival compared to those solely based on baseline information. Coupling multivariate functional principal component analysis (MFPCA) with Cox regression (MFPCA-Cox) provided the highest predictive discrimination and accuracy for the NSCLC patients with AUC values of 0.7 - >0.8 at various landmark time points and prediction timeframes, outperforming recent advanced Transformer and convolutional neural network deep-learning algorithms (TransformerJM and Match-Net, respectively). In conclusion, we identified that MFPCA-Cox represents a robust and versatile joint modeling algorithm for high-dimensional biomarker longitudinal data with irregular and missing data, capturing complex relationships within the data, yielding accurate predictions for both longitudinal biomarkers and survival outcomes, and gaining insights into the underlying dynamics.

Published in BMC Medical Informatics and Decision Making (predicted rank #24) · training set

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.