The Rise of the Large Language Models (LLMs): Can They Truly Match Clinical and Data Science Experts in Clinical Trial Data Analysis?
Komandur Elayavilli, R.; McDunn, J.; Nair, N.; Fall, B.; Dicker, A. P.; Khozin, S.
Show abstract
BackgroundClinical trials provide evidence of the efficacy and safety of experimental treatment regimens. Analysis of data from these trials is a time-intensive process traditionally requiring advanced multidisciplinary expertise in biomedicine, clinical research, biostatistics and data science. Large language models (LLMs), such as OpenAIs GPT-4 and Googles Gemini Advanced, present new opportunities for data analysis in medical research by leveraging natural language understanding and data interpretation capabilities. ObjectiveThis study investigates the ability of LLMs to analyze and report clinical trial results, starting with de-identified individual patient data. Here, we evaluate two LLMs for their ability to recapitulate the analysis of a clinical trial that evaluated LY2510924 in combination with carboplatin and etoposide for the treatment of extensive-stage small cell lung cancer (ES-SCLC). The main objectives are to (i) assess whether LLMs can be effectively used without specialized machine learning training and (ii) compare LLM-driven analyses to those conducted by experienced data scientists. MethodsData from the Project Data Sphere (PDS) platform were used, and multiple investigators employed both ChatGPT and Gemini Advanced for analysis. A chain-of-thought (CoT) prompting framework was applied to guide the LLMs through a systematic evaluation of baseline characteristics, progression-free survival (PFS), overall survival (OS), safety data, and biomarker information. Results were compared across investigators and LLMs to assess consistency. ResultsWhile LLMs could process the trial data and generate relevant insights, discrepancies were observed across the investigators analyses, particularly in primary and secondary endpoints. One investigator found a significant improvement in PFS with LY2510924, contradicting other results. Variations in reporting objective response rate (ORR) and adverse event analyses also highlighted challenges in reporting between different LLMs. These discrepancies may be due to differences in LLM capabilities and behaviors, prompting strategies, and potential model drift over time. ConclusionLLMs such as ChatGPT-4 and Gemini Advanced offer promising capabilities in clinical data analysis, though variability in results underscores the need for tailored CoT frameworks and specialized prompting strategies. Addressing issues such as model drift and ensuring consistent model versions are crucial for reliable application. LLMs also show potential to accelerate clinical research by drafting clinical trial reports, but further refinements are needed to ensure accuracy and consistency in their application. The observed discrepancies across LLM results and in comparison, to the expert-authored trial report highlight the need for highly trained subject matter experts to review and revise LLM-generated clinical trial analyses.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Analysis of clinical trial registry entry histories using the novel R package cthist 94%
- An interactive retrieval system for clinical trial studies with context-dependent protocol elements 94%
- Using numerical modelling and simulation to assess the ethical burden in clinical trials and how it relates to the proportion of responders in a trial sample 94%
Similar papers in this journal
- Supporting Reanalysis and Reuse of Clinical Trial Data: A Case Study 95%
- Controlled evaLuation of Angiotensin Receptor Blockers for COVID-19 respIraTorY disease (CLARITY): Statistical analysis plan for a randomised controlled Bayesian adaptive sample size trial 94%
- Trials that turn from retrospectively registered to prospectively registered: A cohort study of ‘retroactively prospective’ clinical trial registration using history data 92%
Similar papers in this journal
Similar papers in this journal
- Use of unstructured text in prognostic clinical prediction models: a systematic review 94%
- A Web-based Tool for Automatically linking Clinical Trials to their Publications 93%
- Is One Run Enough? Reproducibility of Flagship Large Language Models Across Temperature and Reasoning Settings in Biomedical Text Processing 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.