Deep latent variable modelling reveals clinically significant subgroups among transfusion recipients
Peltola, E.; Turkulainen, E.; Heinonen, M.; Arvas, M.; Ilmakunnas, M.; Koskinen, M.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWO_ST_ABSBackgroundC_ST_ABSTransfusion recipients are a heterogeneous group of patients, yet the identification of these groups has traditionally relied on human-driven univariate analyses and domain knowledge instead of analyzing multivariate characteristics of individuals. Electronic health records (EHR) combined with unsupervised machine learning enables robust, data-driven way for phenotyping patient populations, providing finer-grained view on subgroup characteristics. Materials and MethodsWe introduce an extension to the Variational Autoencoder (VAE) framework and apply the model to EHR data of 19,629 adult transfusion recipients. The latent representation of VAEs approximates a low-dimensional manifold of input data, in which patients with similar characteristics are embedded close to one another. The model integrates clustering via a Gaussian Mixture Model (GMM) prior to identify clinically relevant patient subgroups from diagnosis codes, laboratory values and demographics, while simultaneously classifying the type of transfused products. The final clusters are derived with a modified consensus clustering approach. ResultsWe identified six patient groups with distinct diagnosis, laboratory, demographic, and transfusion profiles. These novel clusters provide a refined characterization of transfusion-related phenotypes, revealing more detailed distinctions among patient subgroups. Our model achieved moderate classification accuracy, with AUROC of 0.879, 0.806 and 0.861, and PR-AUC of 0.448, 0.357 and 0.492 for red blood cells (RBC), plasma and platelets, respectively. Clustering accuracy remains consistent across training and test sets. ConclusionsData-driven phenotyping of transfusion recipients revealed previously unexplored patient phenotypes differing in multiple characteristics. The model helps to understand the heterogeneous nature of patients requiring transfusion and provides insights on how different blood product profiles shift cluster assignments. These findings underscore the utility of latent variable modelling for population characterization and suggest potential applications in optimizing transfusion strategies as well as blood supply chain management. Validation in external cohorts remains unestablished.
Matching journals
The top 12 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Longitudinal laboratory testing tied to PCR diagnostics in COVID-19 patients reveals temporal evolution of distinctive coagulopathy signatures 91%
- An agnostic study of associations between ABO and RhD blood group and phenome-wide disease risk 91%
- Dynamic Tracking of Native Precursors in Adult Mice 91%
Similar papers in this journal
- Machine Learning for Dynamic and Short-term Prediction of Preeclampsia Using Routine Clinical and Laboratory Data 88%
- Diagnostic Codes in AI prediction models and Label Leakage of Same-admission Clinical Outcomes 88%
- Epidemiology and costs of post-sepsis morbidity, nursing care dependency, and mortality in Germany 87%
Similar papers in this journal
- Array Genotyping of Transfusion Relevant Blood Cell Antigens in 6946 Ancestrally Diverse Subjects 92%
- Platelet dysfunction reversal with cold-stored vs. room temperature-stored platelet transfusions 91%
- Hematopoietic recovery after transplantation is primarily derived from the stochastic contribution of hematopoietic stem cells 91%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.