Back

LabQAR: A Manually Curated Dataset for Question Answering on Laboratory Test Reference Ranges and Interpretation

Bhasuran, B.; Jin, Q.; Deville, A.; Wu, Y.; Hanna, K.; Lu, Z.; He, Z.

2025-06-03 health informatics
10.1101/2025.06.03.25328882 medRxiv
Show abstract

Laboratory tests are crucial for diagnosing and managing health conditions, providing essential reference ranges for result interpretation. The diversity of lab tests, influenced by variables like the specimen type (e.g., blood, urine), gender, age-specific, and other influencing factors such as pregnancy, makes automated interpretation challenging. Automated clinical decision support systems attempting to interpret these values must account for such nuances to avoid misdiagnoses or incorrect clinical decisions. In this regard, we present LabQAR (Laboratory Question Answering with Reference Ranges), a manually curated dataset comprising 550 lab test reference ranges derived from authoritative medical sources, encompassing 363 unique lab tests and including multiple-choice questions with annotations on reference ranges, specimen types, and other factors impacting interpretation. We also assess the performance of several large language models (LLMs), including LLaMA 3.1, GatorTronGPT, GPT-3.5, GPT-4, and GPT-4o, in predicting reference ranges and classifying results as normal, low, or high. The findings indicate that GPT-4o outperforms other models, showcasing the potential of LLMs in clinical decision support.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.