Back

Benchmarking Speech Recognition Models for Medical Consultations in Latin American Spanish: A Comparative Evaluation with Fine-Tuning

Carrillo, R. M.; Carbajal Serrano, A.; Condori Pinedo, P. S.

2026-07-16 public and global health
10.64898/2026.07.14.26358062 medRxiv
Show abstract

BACKGROUND: Artificial intelligence (AI) medical scribes rely on speech-to-text (STT) models for transcription. Evaluations of STT models in non-English settings remain scarce. We benchmarked ten STT models on medical consultations from Latin American (LatAm) Spanish and assessed whether fine-tuning improves transcription accuracy. METHODS: Ten YouTube videos depicting medical consultations. Human transcriptions were the ground truth. Five open-source models were evaluated: Whisper Large, Whisper Large v3, Whisper Large v3 Turbo, Voxtral Mini 3B, and Canary 1B v2; and so were five close-source models: gpt-4o-transcribe, gpt-4o-mini-transcribe, gemini-2.5-pro, Eleven Labs, and Assembly AI. Whisper Large v3 was fine-tuned. One video was withheld from training. Performance assessed using Word Error Rate (WER), Character Error Rate (CER), BLEU Score, ROUGE-L, BERT Score, and Semantic Similarity on the one withheld video. RESULTS: None of the fine-tuning iterations outperformed the vanilla Whisper Large v3. With the withheld video, Gemini-2.5-pro was the close-source model with the best performance in four of six metrics. In comparison to the close-source models, the fine-tuned model never outperformed the other models (withheld video); conversely, in comparison to the close-source models, the fine-tuned model showed better performance across metrics, for instance: BLEU score (63% vs to 58% for the second-ranking model), BERT (89% vs to 86%), and semantic similarity (89% vs to 83%), CER (19% vs 20%). CONCLUSIONS: Whisper Large v3 and its fine-tuned variant are the best open-source STT models for transcribing medical conversations in LatAm Spanish. These findings provide an evidence base for developing AI medical scribes tailored to Spanish-speaking LatAm.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

1
PLOS ONE
5266 papers in training set
Top 12%
15.4%
2
Scientific Reports
3612 papers in training set
Top 12%
6.4%
3
Journal of Neural Engineering
221 papers in training set
Top 0.6%
5.6%
4
Frontiers in Digital Health
24 papers in training set
Top 0.3%
4.4%
5
Scientific Data
209 papers in training set
Top 0.6%
4.1%
6
Journal of Medical Internet Research
87 papers in training set
Top 0.7%
3.5%
7
PLOS Digital Health
106 papers in training set
Top 2%
3.3%
8
Biology Methods and Protocols
61 papers in training set
Top 0.3%
3.3%
9
npj Digital Medicine
118 papers in training set
Top 2%
2.7%
10
Frontiers in Psychology
56 papers in training set
Top 0.5%
2.4%
50% of probability mass above
11
JMIRx Med
32 papers in training set
Top 0.9%
1.8%
12
Epidemics
116 papers in training set
Top 1%
1.4%
13
Journal of the American Medical Informatics Association
71 papers in training set
Top 2%
1.4%
14
DIGITAL HEALTH
17 papers in training set
Top 0.6%
1.4%
15
Frontiers in Neuroscience
256 papers in training set
Top 4%
1.4%
16
PLOS Global Public Health
344 papers in training set
Top 7%
1.2%
17
Healthcare
17 papers in training set
Top 0.5%
1.2%
18
Bioinformatics Advances
203 papers in training set
Top 4%
1.2%
19
IEEE Access
35 papers in training set
Top 0.9%
1.2%
20
BMC Medicine
176 papers in training set
Top 3%
1.2%
21
Journal of Speech, Language, and Hearing Research
13 papers in training set
Top 0.1%
1.2%
22
Brain Sciences
55 papers in training set
Top 1%
1.1%
23
Journal of Clinical Medicine
97 papers in training set
Top 4%
1.1%
24
BMJ Health & Care Informatics
15 papers in training set
Top 0.8%
1.0%
25
Journal of Biomedical Informatics
47 papers in training set
Top 1%
0.9%
26
JAMIA Open
42 papers in training set
Top 1%
0.9%
27
Heliyon
152 papers in training set
Top 7%
0.9%
28
Medicine
31 papers in training set
Top 2%
0.9%
29
Expert Systems with Applications
11 papers in training set
Top 0.4%
0.9%
30
npj Genomic Medicine
36 papers in training set
Top 0.7%
0.9%