LabQAR: A Manually Curated Dataset for Question Answering on Laboratory Test Reference Ranges and Interpretation
Bhasuran, B.; Jin, Q.; Deville, A.; Wu, Y.; Hanna, K.; Lu, Z.; He, Z.
Show abstract
Laboratory tests are crucial for diagnosing and managing health conditions, providing essential reference ranges for result interpretation. The diversity of lab tests, influenced by variables like the specimen type (e.g., blood, urine), gender, age-specific, and other influencing factors such as pregnancy, makes automated interpretation challenging. Automated clinical decision support systems attempting to interpret these values must account for such nuances to avoid misdiagnoses or incorrect clinical decisions. In this regard, we present LabQAR (Laboratory Question Answering with Reference Ranges), a manually curated dataset comprising 550 lab test reference ranges derived from authoritative medical sources, encompassing 363 unique lab tests and including multiple-choice questions with annotations on reference ranges, specimen types, and other factors impacting interpretation. We also assess the performance of several large language models (LLMs), including LLaMA 3.1, GatorTronGPT, GPT-3.5, GPT-4, and GPT-4o, in predicting reference ranges and classifying results as normal, low, or high. The findings indicate that GPT-4o outperforms other models, showcasing the potential of LLMs in clinical decision support.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 95%
- Development and Prospective Implementation of a Large Language Model based System for Early Sepsis Prediction 95%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Biomedical Text Normalization through Generative Modeling 94%
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 94%
- A Deep Learning Approach for Transgender and Gender Diverse Patient Identification in Electronic Health Records 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.