Back

Evaluating Large Language Models for Translating Multimodal Phenotype Documentations into Executable EHR Phenotyping Algorithms

Yan, C.; Xin, Y.; Su, W.-C.; Gangireddy, S.; Durbhakula, S.; Bruehl, S. P.; Dickson, A. L.; Li, L.; Feng, Q.; Malin, B. A.; Derr, T.; Wei, W.-Q.

2026-05-22 health informatics
10.64898/2026.05.20.26353690 medRxiv
Show abstract

Research applications of electronic health record (EHR) phenotypes require translating clinical definitions into executable EHR database queries, a labor-intensive process. We evaluated two frontier large language models across five phenotypes and three documentation modalities. Both models captured high-level logic from structured text but degraded markedly with diagram-only input. Error analysis revealed seven failure categories. Documentation, rather than model capability, was the primary bottleneck, reinforcing the need for standardization and expert oversight.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.