Back

Uncertainty-aware extraction of clinical findings from Finnish EHRs using open large language models

Leinonen, J. V.; Knuutila, J.; Kurki, S.; Pamilo, S.; Koskinen, M.

2026-07-09 health informatics
10.64898/2026.07.07.26355248 medRxiv
Show abstract

Objective. To evaluate whether open-weight large language models (LLMs) can accurately extract clinical findings from Finnish-language pediatric records, and whether prediction uncertainty can be used to triage cases for expert review to minimize manual work. Materials and Methods. Retrospective cohort of 97 pediatric ischaemic stroke patients (1 month - 17 years) from Helsinki University Hospital (2010 - 2023). Three open LLMs (gpt-oss-20b, DeepSeek-R1-Distill-Qwen-32B, and medgemma-27b-text-it) were prompted in English to detect four extraction targets (hemiplegia, headache, seizure, and stroke as a positive control) from each patient's full free-text record. Each combination received 15 calls (five temperatures x three repeats). Performance was benchmarked against a clinician reference (accuracy, recall, precision, F1). Shannon entropy across the 15 calls quantified within-model uncertainty; inter-model disagreement provided an ensemble signal. Patients were ranked by uncertainty for a simulated selective-review workflow. Findings were externally validated in an independent neonatal stroke cohort (n = 88). Results. Gpt-oss-20b achieved the best balance of recall (0.91 - 1.00) and precision (0.83 - 0.92), with F1 0.89 - 0.95 across non-control extraction targets. Entropy in misclassified cases was 2.4 - 3.4 times higher than in correctly classified cases. Entropy-based triage achieved complete error coverage by reviewing <10% of patients for hemiplegia (8.3%) and headache (8.2%), and 19.6% for seizure. Neonatal validation reached F1 0.95 for Apgar 1 min and binary seizure, and F1 0.87 for 4-class stroke-subtype classification. Discussion. Within-model entropy and inter-model disagreement provided complementary, calibrated signals of likely error in a non-English clinical setting. Conclusion. Open LLMs can extract clinical findings from Finnish pediatric records with accuracy comparable to published English benchmarks, and uncertainty-based triage substantially reduces required expert workload.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
Journal of the American Medical Informatics Association
71 papers in training set
Top 0.2%
18.0%
2
BMC Medical Informatics and Decision Making
43 papers in training set
Top 0.2%
7.7%
3
Scientific Reports
3612 papers in training set
Top 12%
6.5%
4
BMJ Health & Care Informatics
15 papers in training set
Top 0.1%
6.1%
5
JMIR Medical Informatics
18 papers in training set
Top 0.1%
5.3%
6
PLOS ONE
5266 papers in training set
Top 30%
5.0%
7
npj Digital Medicine
118 papers in training set
Top 1%
5.0%
50% of probability mass above
8
PLOS Digital Health
106 papers in training set
Top 1%
4.7%
9
Frontiers in Digital Health
24 papers in training set
Top 0.2%
4.7%
10
Journal of Biomedical Informatics
47 papers in training set
Top 0.4%
3.9%
11
BMC Medical Research Methodology
47 papers in training set
Top 0.6%
1.9%
12
iScience
1154 papers in training set
Top 18%
1.7%
13
Communications Medicine
113 papers in training set
Top 2%
1.7%
14
BMJ
51 papers in training set
Top 0.6%
1.6%
15
eBioMedicine
183 papers in training set
Top 3%
1.6%
16
JAMA Pediatrics
10 papers in training set
Top 0.1%
1.5%
17
The Lancet Digital Health
25 papers in training set
Top 0.4%
1.3%
18
International Journal of Medical Informatics
26 papers in training set
Top 1%
1.1%
19
Computer Methods and Programs in Biomedicine
28 papers in training set
Top 0.8%
1.1%
20
Artificial Intelligence in Medicine
17 papers in training set
Top 0.5%
1.1%
21
JAMA Network Open
130 papers in training set
Top 3%
1.0%
22
European Heart Journal - Digital Health
18 papers in training set
Top 1%
0.8%
23
Nature Communications
5641 papers in training set
Top 58%
0.8%
24
Journal of Clinical Epidemiology
31 papers in training set
Top 0.8%
0.8%
25
JAMIA Open
42 papers in training set
Top 2%
0.8%
26
Genetics in Medicine
78 papers in training set
Top 1%
0.8%
27
Orphanet Journal of Rare Diseases
21 papers in training set
Top 0.7%
0.6%
28
Genome Medicine
183 papers in training set
Top 6%
0.6%
29
BMJ Open
601 papers in training set
Top 14%
0.6%