Back

Fine-tuning and pre-training improve the predictive accuracy of large language models for rheumatoid arthritis disease activity

Honda, S.; Ikari, K.; Fujisaki, M.; Tanaka, E.; Harigai, M.

2024-10-16 rheumatology
10.1101/2024.10.14.24315448 medRxiv
Show abstract

ObjectiveTo evaluate whether the performance of the large language model (LLM) Llama2 improves with pre-training and fine-tuning, and to compare its predictive accuracy with that of a linear regression model for rheumatoid arthritis (RA) disease activity. MethodsClinical data from 11,865 patients in the cohort were used to predict disease activity at two years on four indices (Disease Activity Score (DAS) 28-Erythrocyte sedimentation rate (ESR), DAS28-C-reactive protein (CRP), Clinical Disease Activity Index (CDAI) or Japanese Health Assessment Questionnaire (J-HAQ)). Logistic regression was employed for the linear model for comparison. The predictive performance was assessed using area under the curve (AUC). Additional performance metrics including precision, recall, and F1 score were calculated. ResultsPre-training significantly improved AUC of Meditron (Llama2 pre-trained with medical data) for DAS28-ESR >5.1, DAS28-ESR <2.6, DAS28-CRP <2.3, J-HAQ score >2.5, and J-HAQ score <0.5 (P <0.05). Fine-tuning resulted in significant improvements in AUC for Llama2 across all indices (P <0.05) except CDAI >22, and for Meditron in DAS28-ESR <2.6, DAS28-CRP >4.1, DAS28-CRP <2.3 and CDAI [&le;]2.8 (P <0.05). Both LLMs significantly outperformed linear regression in predicting DAS28-ESR <2.6, DAS28-CRP >4.1, DAS28-CRP <2.3, J-HAQ score >2.5, and J-HAQ score <0.5 (P <0.05). Furthermore, DAS28-CRP >4.1, DAS28-CRP <2.3, J-HAQ score >2.5 and J-HAQ score <0.5, Llama2 or Meditron consistently outperformed linear regression across all performance metrics. ConclusionBoth pre-training and fine-tuning significantly improved the performance of Llama2. Both LLMs outperformed the linear regression model in predicting 5 out of the 8 categories of indices.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.