Towards Maps of Disease Progression: Biomedical Large Language Model Latent Spaces For Representing Disease Phenotypes And Pseudotime
Zamora-Resendiz, R.; Khurram, I.; Crivelli, S.
Show abstract
In this study, we propose a scientific framework to detect capability among biomedical large language models (LLMs) for organizing expressions of comorbid disease and temporal progression. We hypothesize that biomedical LLMs pretrained on next-token prediction produce latent spaces that implicitly capture "disease states" and disease progression, i.e., the transitions over disease states over time. We describe how foundation models may capture and transfer knowledge from explicit pretraining tasks to specific clinical applications. A scoring function based on Kullback-Leibler divergence was developed to measure "surprise" in seeing specialization when subsetting admissions along 13 biomedical LLM latent spaces. By detecting implicit ordering of longitudinal data, we aim to understand how these models self-organize clinical information and support tasks such as phenotypic classification and mortality prediction. We test our hypothesis along a case study for obstructive sleep apnea (OSA) in the publicly available MIMIC-IV dataset, finding ordering of phenotypic clusters and temporality within latent spaces. Our quantitative findings suggest that increased compute, conformance with compute-optimal training, and widening contexts promote better implicit ordering of clinical admissions by disease states, explaining 60.3% of the variance in our proposed implicit task. Preliminary qualitative findings suggest LLMs latent spaces trace patient trajectories through different phenotypic clusters, terminating at end-of-life phenotypes. This approach highlights the potential of biomedical LLMs in modeling disease progression, identifying new patterns in disease pathways and interventions, and evaluating clinical hypotheses related to drivers of severe illness. We underscore the need for larger, high-resolution longitudinal datasets to further validate and enhance understanding of the utility of LLMs in modeling patient trajectories along clinical text and advancing precision medicine. Key PointsO_ST_ABSQuestionC_ST_ABSDo LLMs sensibly organize cilnical data with respect to applications in precision medicine? FindingsBiomedically-trained LLMs show increasing potential in promoting the organization of patient data to reflect disease progression. In a subcohort of OSA patients, maps derived from LLMs latent representations reveal traceable disease trajectories. MeaningMaps of disease progression offer an explanation to the utility of LLMs in precision medicine. Following current pretraining conventions in foundation modeling, scientific inquiry into these maps may help anticipate progress in applications of LLMs for healthcare.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Continuous-Time and Dynamic Suicide Attempt Risk Prediction with Neural Ordinary Differential Equations 95%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 95%
- Zero Shot Health Trajectory Prediction Using Transformer 95%
Similar papers in this journal
- ARCH: Large-scale Knowledge Graph via Aggregated Narrative Codified Health Records Analysis 95%
- A scoping review of fair machine learning techniques when using real-world data 94%
- A methodology of phenotyping ICU patients from EHR data: high-fidelity, personalized, and interpretable phenotypes estimation 94%
Similar papers in this journal
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 95%
- EHR Foundation Models Improve Robustness in the Presence of Temporal Distribution Shift 95%
- Using explainable machine learning to identify patients at risk of reattendance at discharge from emergency departments 94%
Similar papers in this journal
- Comparing neural language models for medical concept representation and patient trajectory prediction 95%
- Enriching Representation Learning Using 53 Million Patient Notes through Human Phenotype Ontology Embedding 94%
- Deep ensemble multitask classification of emergency medical call incidents combining multimodal data improves emergency medical dispatch 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.