An In-Depth Evaluation of Federated Learning on Biomedical Natural Language Processing
Peng, L.; Luo, G.; Zhou, S.; Chen, J.; Xu, Z.; Zhang, R.; Sun, J.
Show abstract
Language models (LMs) such as BERT and GPT have revolutionized natural language processing (NLP). However, the medical field faces challenges in training LMs due to limited data access and privacy constraints imposed by regulations like the Health Insurance Portability and Accountability Act (HIPPA) and the General Data Protection Regulation (GDPR). Federated learning (FL) offers a decentralized solution that enables collaborative learning while ensuring data privacy. In this study, we evaluated FL on 2 biomedical NLP tasks encompassing 8 corpora using 6 LMs. Our results show that: 1) FL models consistently outperformed models trained on individual clients data and sometimes performed comparably with models trained with polled data; 2) with the fixed number of total data, FL models training with more clients produced inferior performance but pre-trained transformer-based models exhibited great resilience. 3) FL models significantly outperformed large language models using zero-/one-shot learning and offered lightning inference speed.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Federated Learning for multi-omics: a performance evaluation in Parkinson's disease 95%
- Inferring global-scale temporal latent topics from news reports to predict public health interventions for COVID-19 95%
- Building a Best-in-Class De-identification Tool for Electronic Medical Records Through Ensemble Learning 94%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.