Back

Clinical Note Comparison and Data Retrieval Via Embedding Vectors: Model Selection, Metrics, and Convergence

Dahlberg, A. C. H.; Tapiola, O.; Luisto, R.; Puranen, T.; Sanmark, E.; Vartiainen, V.

2026-05-18 health informatics
10.64898/2026.05.12.26352832 medRxiv
Show abstract

Background: Embedding models are an integral part of generative AI architectures, transforming text into embedding vectors that represent semantic content in numerical form. Despite their central role, their performance in clinical settings remains underexplored. We evaluate embedding models across two tasks: semantic difference detection in clinical texts, and data retrieval from patient records. Methods: Eight models were applied to synthetic discharge summaries in English, Finnish, and Swedish. Semantic sensitivity was assessed by introducing controlled perturbations (deletion, modification, and paraphrasing) at three levels of severity; cosine similarity, and L1 and Euclidean distances were computed between the vectors of the original and perturbed texts. Partial vectors were compared to explore dimensionality reduction. Two models with the biggest contrast in semantic difference detection were evaluated on retrieval of relevant information from real Finnish vascular surgery records. Results: Embedding vectors captured semantic differences in clinical text: content deletion and modification produced larger increases in vector distance than paraphrasing. On average, models detected the direction of semantic change correctly, but case-level performance varied considerably. Qwen3-Embedding-8B was the only model with zero directional errors, while multilingual-E5-large erred in 13.8% of cases. In data retrieval, Qwen3-Embedding-8B again outperformed multilingual-E5-large, though the margin was narrower: sufficiency scores were 3.25 vs. 3.17 out of 5 for the first query and 2.25 vs. 1.15 out of 5 for the second query. For some models, as few as 0.6-1.2% of dimensions sufficed to replicate full-vector accuracy; principal component analysis and coordinate-level analysis did not account for this finding. Conclusions: Our results show that the choice of embedding model is important: performance differences between models can be large enough to determine whether clinically relevant information reaches the end user, and model weaknesses can be both task-specific and context-dependent.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
npj Digital Medicine
118 papers in training set
Top 0.5%
12.5%
2
Artificial Intelligence in Medicine
17 papers in training set
Top 0.1%
9.7%
3
Journal of the American Medical Informatics Association
71 papers in training set
Top 0.3%
9.7%
4
BMJ Health & Care Informatics
15 papers in training set
Top 0.1%
6.6%
5
BMC Medical Informatics and Decision Making
43 papers in training set
Top 0.3%
6.6%
6
JMIR Medical Informatics
18 papers in training set
Top 0.1%
4.0%
7
Scientific Reports
3612 papers in training set
Top 27%
4.0%
50% of probability mass above
8
Journal of Biomedical Informatics
47 papers in training set
Top 0.4%
4.0%
9
Frontiers in Digital Health
24 papers in training set
Top 0.3%
3.5%
10
Journal of Medical Internet Research
87 papers in training set
Top 0.7%
3.4%
11
PLOS Digital Health
106 papers in training set
Top 2%
3.2%
12
PLOS ONE
5266 papers in training set
Top 39%
3.1%
13
Frontiers in Artificial Intelligence
20 papers in training set
Top 0.2%
2.4%
14
Biology Methods and Protocols
61 papers in training set
Top 0.5%
2.4%
15
International Journal of Medical Informatics
26 papers in training set
Top 0.6%
2.1%
16
Computers in Biology and Medicine
128 papers in training set
Top 2%
1.9%
17
JAMIA Open
42 papers in training set
Top 0.8%
1.9%
18
Patterns
78 papers in training set
Top 2%
1.5%
19
IEEE Journal of Biomedical and Health Informatics
37 papers in training set
Top 0.9%
1.1%
20
Journal of Neural Engineering
221 papers in training set
Top 2%
1.1%
21
BMC Bioinformatics
457 papers in training set
Top 5%
1.0%
22
iScience
1154 papers in training set
Top 32%
0.9%
23
DIGITAL HEALTH
17 papers in training set
Top 0.9%
0.8%
24
Medicine
31 papers in training set
Top 2%
0.8%
25
Computer Methods and Programs in Biomedicine
28 papers in training set
Top 1%
0.8%