Back

Combining Clinician Expertise with Prompt Engineering enhances Small Language Models Reliability for Cancer Entity Recognition in Electronic Health Records

Corso, F.; Peppoloni, V.; Mazzeo, L.; Leone, G.; Passos, L.; Miskovic, V.; Armanini, J.; Ferrarin, A.; Wiest, I. C.; Wolf, F.; Montelatici, G.; Romano', R.; Ambrosini, P.; Capoccia, T.; Natangelo, S.; Rota, S.; Andena, P.; De Ponti, M.; Russo, A.; Stasi, G.; Provenzano, L.; Spagnoletti, A.; Meazza Prina, M.; Cavalli, C.; Giani, C.; Serino, R.; Borraccino, M.; Bonalume, C.; Di Mauro, R. M.; Agosta, C.; Dumitrascu, A. D.; Di Liberti, G.; Corrao, G.; Beninato, T.; Ganzinelli, M.; Occhipinti, M.; Brambilla, M.; Proto, C.; Kather, J. N.; Pedrocchi, A. L. G.; De Braud, F.; Lo Russo, G.; Baili, P.; P

2025-10-21 oncology
10.1101/2025.10.16.25337917 medRxiv
Show abstract

Real-world data (RWD), largely stored in unstructured electronic health records (EHRs), are critical for understanding complex diseases like cancer. However, extracting structured information from these narratives is challenging due to linguistic variability, semantic complexity, and privacy concerns. This study evaluates the performance of four locally deployable and small language models (SLMs), LLaMA, Mistral, BioMistral, and MedLLaMA, for information extraction (IE) from Italian EHRs within the APOLLO 11 trial on non-small cell lung cancer (NSCLC). We examined three prompting strategies (zero-shot, few-shot, and annotated few-shot) across English and Italian, involving clinicians with varying expertise to assess prompt designs impact on accuracy. Results show that general-purpose models (e.g., LLaMA 3.1 8B) outperform biomedical models in most tasks, particularly in extracting binary features. Multiclass variables such as TNM staging, PD-L1, and ECOG were more difficult due to implicit language and lack of standardization. Few-shot prompting and native-language inputs significantly improved performance and reduced hallucinations. Clinical expertise enhanced consistency in annotation, particularly among students using annotated examples. The study confirms that privacy-preserving SLMs can be deployed locally for efficient and secure cancer data extraction. Findings highlight the need for hybrid systems combining SLMs with expert input and underline the importance of aligning clinical documentation practices with SLM capabilities. This is the first study to benchmark SLMs on Italian EHRs and investigate the role of clinical expertise in prompt engineering, offering valuable insights for the future integration of SLMs into real-world clinical workflows.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
npj Digital Medicine
118 papers in training set
Top 0.2%
22.4%
2
Artificial Intelligence in Medicine
17 papers in training set
Top 0.1%
12.9%
3
Scientific Reports
3612 papers in training set
Top 10%
6.9%
4
Journal of the American Medical Informatics Association
71 papers in training set
Top 0.6%
5.6%
5
Communications Medicine
113 papers in training set
Top 0.7%
4.1%
50% of probability mass above
6
JCO Clinical Cancer Informatics
22 papers in training set
Top 0.2%
3.6%
7
PLOS ONE
5266 papers in training set
Top 35%
3.6%
8
Database
61 papers in training set
Top 0.2%
3.3%
9
JAMIA Open
42 papers in training set
Top 0.6%
2.7%
10
JMIR Medical Informatics
18 papers in training set
Top 0.3%
2.5%
11
iScience
1154 papers in training set
Top 10%
2.5%
12
Diagnostics
50 papers in training set
Top 0.8%
2.4%
13
Biology Methods and Protocols
61 papers in training set
Top 0.5%
2.2%
14
eBioMedicine
183 papers in training set
Top 2%
1.7%
15
Frontiers in Digital Health
24 papers in training set
Top 0.8%
1.5%
16
JAMA Network Open
130 papers in training set
Top 2%
1.5%
17
PLOS Digital Health
106 papers in training set
Top 3%
1.1%
18
Journal of Medical Internet Research
87 papers in training set
Top 2%
1.1%
19
Journal of Biomedical Informatics
47 papers in training set
Top 1%
1.1%
20
Computers in Biology and Medicine
128 papers in training set
Top 4%
1.0%
21
Nature Communications
5641 papers in training set
Top 55%
0.9%
22
Computer Methods and Programs in Biomedicine
28 papers in training set
Top 1.0%
0.9%
23
DIGITAL HEALTH
17 papers in training set
Top 0.9%
0.9%
24
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 40%
0.9%
25
BMC Medical Research Methodology
47 papers in training set
Top 1%
0.9%
26
PeerJ
308 papers in training set
Top 12%
0.6%