Predictive and Explainable Analysis of Post-operative Acute Kidney Injury in Children undergoing Cardiopulmonary Bypass: An Application of Large Language Models
Sharabiani, M. T. A.; Mahani, A. S.; Bottle, A.; Srinivasan, Y.; Issitt, R. W.; Stoica, S.
Show abstract
The emergence of large language models (LLMs) offers new opportunities to leverage, often unused, information in clinical text. This study examines the utility of text embeddings generated by LLMs in predicting postoperative acute kidney injury (AKI) in paediatric cardiopulmonary bypass (CPB) patients using electronic health record (EHR) text, and to explore methods for explaining their output. AKI is a significant complication in paediatric CPB and its prediction can significantly improve patient outcomes by enabling timely interventions. We evaluate various text embedding algorithms such as Doc2Vec, top-performing sentence transformers on Hugging Face, and commercial LLMs from Google and OpenAI. We benchmark the out-of-sample predictive performance of these AI models against a baseline model as well as an established clinically-defined expert model. The baseline model includes patient gender, age, height, body mass index and length of operation. The majority of AI models surpass, not only the baseline model, but also the expert model. An ensemble of AI and clinical-expert models improves discriminative performance by nearly 23% compared to the baseline model. Consistency of patient clusters formed from AI-generated embeddings with clinical-expert clusters - measured via the adjusted rand index and adjusted mutual information metrics - illustrates their medical validity. We use text-generating LLMs to explain the output of embedding LLMs, e.g., by summarising the differences between AI and expert clusters, and/or by providing descriptive labels for the AI clusters. Such explainability can increase medical practitioners trust in the AI applications, and help generate new hypotheses, e.g., by correlating cluster memberships with outcomes of interest. HighlightsO_LILLMs outperform clinical experts in predicting risk of AKI after paediatric CPB. C_LIO_LILLMs generate clinically plausible explanations and hypotheses using embeddings. C_LIO_LISuccessful application of LLMs in paediatric CPB suggests potential in other specialised fields. C_LIO_LIFine-tuning LLMs on domain data and forming ensembles of AI and clinical experts may boost accuracy. C_LI
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 95%
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 94%
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 94%
Similar papers in this journal
- Comparing neural language models for medical concept representation and patient trajectory prediction 97%
- Enriching Representation Learning Using 53 Million Patient Notes through Human Phenotype Ontology Embedding 96%
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 94%
Similar papers in this journal
Similar papers in this journal
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 96%
- Explainable deep learning for disease activity prediction in chronic inflammatory joint diseases 94%
- Evaluating Knowledge Fusion Models on Detecting Adverse Drug Events in Text 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.