Forecasting Alzheimer's Disease Progression with Deep Multimodal Learning: Integration of 3D MRI and Tabular Clinical Records via a Large Vision-Language Model
Dang, P.; Yu, F.; Guo, T.; Yan, J.; Zhang, C.; Cao, S.; Hashemifar, S.
Show abstract
BackgroundAccurate forecasting of Alzheimers Disease (AD) progression is critical for personalized patient management and clinical trial stratification. However, current predictive models often struggle to effectively integrate high-dimensional neuroimaging with longitudinal clinical data. We introduce AD-LLaVA-3D, a novel multimodal framework designed to bridge this gap by adapting large vision-language models for volumetric and temporal forecasting. MethodsWe leveraged the LLaVA-NeXT-Video architecture to treat 3D MRI volumes as temporal sequences, enabling the model to process volumetric imaging alongside longitudinal Tabular Clinical Records (TCR). The model was trained on the Alzheimers Disease Neuroimaging Initiative (ADNI) cohort (n=764) and evaluated using a rigorous patient-level split. We assessed its ability to forecast a suite of future clinical indicators (e.g., CDR-SB, MMSE) against traditional machine learning baselines (Lasso, Random Forest, Gradient Boosting) and specialized deep learning models (ResNet-3D, Med-Flamingo). ResultsAD-LLaVA-3D demonstrated superior predictive accuracy on the ADNI test set, achieving a Coefficient of Determination (R2) of 0.68 for the critical CDR-SB score, surpassing the best-performing baseline (R2 = 0.66). Crucially, in an independent external validation on the Open Access Series of Imaging Studies (OASIS) cohort (n=76), our model exhibited exceptional generalization (R2 = 0.82, MSE = 0.54), whereas comparison models showed significant performance degradation (R2 < 0.60). ConclusionsThis study presents the first application of a video-based multimodal architecture for AD progression forecasting. By effectively integrating 3D MRI with tabular clinical records, AD-LLaVA-3D offers a robust, generalizable tool for monitoring disease trajectories, significantly advancing predictive capabilities beyond current unimodal or static methods. HighlightsFirst-in-Class Architecture: We introduce the first application of video-based Large Vision-Language Models (LVLMs) to interpret 3D volumetric MRI as a temporal sequence, capturing longitudinal neurodegeneration more effectively than static 3D-CNNs. Robust External Validation: The model achieved superior predictive accuracy (R2 = 0.82) on an independent external cohort (OASIS), demonstrating exceptional generalization beyond the training population (ADNI). Data-Efficient Multimodal Integration: We developed a novel prompting strategy that integrates sparse Tabular Clinical Records (TCR) without artificial imputation, allowing the model to leverage incomplete real-world medical history. Clinical Trial Enrichment: By accurately forecasting future cognitive scores (CDR-SB, MMSE), AD-LLaVA-3D serves as a precise screening tool to identify "rapid progressors" for clinical trials, potentially reducing failure rates in drug development.1
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Clinical Neuroimaging Platform for Rapid, Automated Lesion Detection and Personalized Post-Stroke Outcome Prediction 94%
- Interpretable deep learning approach for extracting cognitive features from hand-drawn images of intersecting pentagons in older adults 93%
- Biologically-informed deep neural networks provide quantitative assessment of intratumoral heterogeneity in post-treatment glioblastoma 92%
Similar papers in this journal
- WMH-DualTasker: A weakly-supervised deep learning model for automated white matter hyperintensities segmentation and visual rating prediction 97%
- Cross-dataset Evaluation of Dementia Longitudinal Progression Prediction Models 96%
- OpenMAP-T1: A Rapid Deep Learning Approach to Parcellate 280 Anatomical Regions to Cover the Whole Brain 95%
Similar papers in this journal
- An explainable self-attention deep neural network for detecting mild cognitive impairment using multi-inputbdigital drawing tasks 94%
- Enhancing MR imaging driven Alzheimers disease classification performance using generative adversarial learning 94%
- ADataViewer: Exploring Semantically Harmonized Alzheimer’s Disease Cohort Datasets 93%
Similar papers in this journal
Similar papers in this journal
- Enhancing Fairness in Disease Prediction by Optimizing Multiple Domain Adversarial Networks 95%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 92%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.