CMANet: Cross-Modal Attention Network for 3-D Knee MRI and Report-Guided Osteoarthritis Assessment
Harari, R. e.; Rajabzadeh-Oghaz, H.; Hosseini, F.; shali, m. g.; Altaweel, A.; Haouchine, N.; Rikhtegar Nezami, F.
Show abstract
ObjectiveKnee osteoarthritis (OA) is a leading cause of disability worldwide, with early identification of structural changes critical for improving patient outcomes. While magnetic resonance imaging (MRI) provides rich spatial detail, its interpretation remains challenging due to complex anatomy, subtle lesion presentation, and limited voxel-level annotations. Meanwhile, radiology reports encode semantic and diagnostic insights that are typically underutilized in imaging AI pipelines. In this work, we introduce CMANet, a Cross-Modal Attention Network that integrates 3D knee MRI volumes with their corresponding free-text radiology reports for joint OA severity classification and lesion segmentation. CMANet introduces four key innovations: (1) an asymmetric cross-modal attention mechanism that enables bidirectional information flow between image and text, (2) a weakly supervised anatomical alignment module linking report phrases to MRI regions, (3) a multi-task prediction head for simultaneous OA grading and voxel-level lesion detection, and (4) interpretable attention pathways for tracing predictions to report language and anatomical structures. Evaluated on a dataset of 642 patients with paired MRI and radiology reports, CMANet achieved significant improvements over unimodal baselines--boosting KL-grade classification AUC from 0.769 to 0.871 ({Delta} =0.102, p=0.004) and increasing Dice scores for cartilage and BML lesion segmentation. The model also demonstrated generalizability in predicting 2-year OA progression (AUC=0.804) and achieved improved alignment between anatomical regions and textual descriptions. These results highlight the potential of multimodal learning to enhance diagnostic accuracy, spatial localization, and explainability in musculoskeletal imaging.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A Clinical Neuroimaging Platform for Rapid, Automated Lesion Detection and Personalized Post-Stroke Outcome Prediction 93%
- Understanding the robustness of vision-language models to medical image artefacts 92%
- STPath: A Generative Foundation Model for Integrating Spatial Transcriptomics and Whole Slide Images 91%
Similar papers in this journal
Similar papers in this journal
- Dual Adversarial Deconfounding Autoencoder for joint batch-effects removal from multi-center and multi-scanner radiomics data 94%
- A multimodal computational pipeline for 3D histology of the human brain 93%
- Improving Rectal Tumor Segmentation with Anomaly Fusion Derived from Anatomical Inpainting: A Multicenter Study 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.