Optimizing an LLM-Based Clinical Data Querying System Using Metadata Enrichment and Task Decomposition
Liu, W.; Qu, B.; Mallya, P.; Wu, J.; Thomas, K.; Hall, J. L.; Zhao, J.; Yin, Z.
Show abstract
Accessing complex clinical registries traditionally requires SQL programming expertise, limiting data accessibility for non-technical researchers. In this paper, we designed and evaluated whether a text-to-SQL solution based on large language models (LLMs) could enable natural language querying of a real-world clinical registry under strict privacy and security constraints. Using self-hosted, open-source LLMs, we developed a multi-layered optimization framework incorporating metadata enrichment, query decomposition, hybrid retrieval, and SQL self-correction. We assessed its performance across 600 queries spanning one-, two-, and three-field complexity using execution-based validation. Accuracy was improved from 88.0% to 94.5% for one-field queries and from 10.0% to 82.0% for three-field queries. Real-world testing by data scientists revealed domain-specific challenges related to coded variables, clinical ambiguity, and multi-step reasoning. We summarize key technical and operational lessons learned and discuss implications for safe, scalable deployment of LLM-assisted analytic tools in clinical registry environments.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Medication information extraction using local large language models 95%
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 94%
- Creating a computer assisted ICD coding system: performance metric choice and use of the ICD hierarchy 94%
Similar papers in this journal
- Comparative Effectiveness of Medical Concept Embedding for Feature Engineering in Phenotyping 94%
- Understanding Data Differences across the ENACT Federated Research Network 93%
- MMFP-Tableau: Enabling Precision Mitochondrial Medicine through Integration, Visualization, and Analytics of Clinical and Research Health System Electronic Data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.