Accelerating Exploratory Clinical Research: An LLM-Powered Framework for Cross-Study Data Harmonization and Natural Language Querying
Garg, A.; Sett, A.; Baumann, B.; Fry, T.; Hedge, S.; Kapadia, B.; Pandit, Y.
Show abstract
Clinical research depends on high quality data that is standardized, accessible and interoperable. Yet evolving data standards over time and variations in their implementation hinder the secondary use of clinical trial datasets. Although individual studies adhere to clinical data standards set forth by CDISC (Clinical Data Interchange Standards Consortium), differences in study design, interpretation of complex models, controlled terminologies, and historical conventions create inconsistencies that limit interoperability and complicate cross-study analysis. As a result, harmonizing SDTM datasets across studies is essential to enable efficient secondary use and accelerate evidence generation. To address these challenges, we introduce a framework that leverages Large Language Models (LLMs) to automate the harmonization of study-specific clinical trial data, available in CDISC Study Data Tabulation Model (SDTM) format, into data that is harmonized across trials at scale and also enables natural language querying via a text-to-SQL agent. This system transforms siloed study clinical datasets into interoperable, analysis-ready formats while empowering users to retrieve insights across trials without needing SQL or domain-specific schema knowledge. By constructing a semantic layer and applying retrieval-augmented prompting to models like GPT-4o, our approach improves data access, query accuracy, and scalability across use cases. This work demonstrates the potential of LLMs to transform clinical data workflows, cutting manual effort, substantially reducing manual harmonization effort and query latency in secondary analysis workflows, and enabling faster exploratory analysis and hypothesis generation in clinical research.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 94%
- MelAnalyze: Fact-Checking Melatonin claims using Large Language Models and Natural Language Inference 92%
- Towards semantic interoperability: finding and repairing hidden contradictions in biomedical ontologies 92%
Similar papers in this journal
- De-novo FAIRification via an Electronic Data Capture system by automated transformation of filled electronic Case Report Forms into machine-readable data 96%
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 96%
- Medication information extraction using local large language models 94%
Similar papers in this journal
- FHIR-DHP: A Standardized Clinical Data Harmonisation Pipeline for scalable AI application deployment 96%
- MIMIC in the OMOP Common Data Model 94%
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.