Forecasting Clinical Risk from Textual Time Series: Structuring Narratives for Temporal AI in Healthcare
Noroozizadeh, S.; Kumar, S.; Weiss, J.
Show abstract
Clinical case reports encode temporal patient trajectories that are often underexploited by traditional machine learning methods relying on structured data. In this work, we introduce the forecasting problem from textual time series, where timestamped clinical findings-extracted via an LLM-assisted annotation pipeline-serve as the primary input for prediction. We systematically evaluate a diverse suite of models, including fine-tuned decoder-based large language models and encoder-based transformers, on tasks of event occurrence prediction, temporal ordering, and survival analysis. Our experiments reveal that encoder-based models consistently achieve higher F1 scores and superior temporal concordance for short- and long-horizon event forecasting, while fine-tuned masking approaches enhance ranking performance. In contrast, instruction-tuned decoder models demonstrate a relative advantage in survival analysis, especially in early prognosis settings. Our sensitivity analyses further demonstrate the importance of time ordering, which requires clinical time series construction, as compared to text ordering, the format of the text inputs that LLMs are classically trained on. This highlights the additional benefit that can be ascertained from time-ordered corpora, with implications for temporal tasks in the era of widespread LLM use.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Building a Best-in-Class De-identification Tool for Electronic Medical Records Through Ensemble Learning 95%
- Inferring global-scale temporal latent topics from news reports to predict public health interventions for COVID-19 94%
- Generating hard-to-obtain information from easy-to-obtain information: applications in drug discovery and clinical inference 93%
Similar papers in this journal
- Generalized Radiograph Representation Learning via Cross-supervision between Images and Free-text Radiology Reports 93%
- COSIME: Cooperative multi-view integration with Scalable and Interpretable Model Explainer 93%
- Estimating Treatment Effects for Time-to-Treatment Antibiotic Stewardship in Sepsis 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.