The Forgotten Shield: Safety Grafting in Parameter-Space for Medical MLLMs
Zhao, J.; Mou, X.; Wu, J.; Yu, H.; Sun, M.; Shi, Y.; Yin, X.; Chen, Z.; Lei, Z.; Wang, Y.
Show abstract
Medical Multimodal Large Language Models (Medical MLLMs) have achieved remarkable progress in specialized medical tasks; however, research into their safety has lagged, posing potential risks for real-world deployment. In this paper, we first establish a multidimensional evaluation framework to systematically benchmark the safety of current SOTA Medical MLLMs. Our empirical analysis reveals pervasive vulnerabilities across both general and medical-specific safety dimensions in existing models, particularly highlighting their fragility against cross-modality jailbreak attacks. Furthermore, we find that the medical fine-tuning process frequently induces catastrophic forgetting of the models original safety alignment. To address this challenge, we propose a novel "Parameter-Space Intervention" approach for efficient safety re-alignment. This method extracts intrinsic safety knowledge representations from original base models and concurrently injects them into the target model during the construction of medical capabilities. Additionally, we design a fine-grained parameter search algorithm to achieve an optimal trade-off between safety and medical performance. Experimental results demonstrate that our approach significantly bolsters the safety guardrails of Medical MLLMs without relying on additional domain-specific safety data, while minimizing degradation to core medical performance.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Compressive Big Data Analytics: An Ensemble Meta-Algorithm for High-dimensional Multisource Datasets 94%
- FedNolowe: A Normalized Loss-Based Weighted Aggregation Strategy for Robust Federated Learning in Heterogeneous Environments 94%
- Deep learning models for COVID-19 chest x-ray classification: Preventing shortcut learning using feature disentanglement 93%
Similar papers in this journal
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 94%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 94%
- Evaluating Knowledge Fusion Models on Detecting Adverse Drug Events in Text 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.