Back

Identification of Patient Trajectories in Timeseries Clinical Transcriptomics Data

Mathur, S.; Chen, Y.; Passban, P.; Zhang, L.; Habiel, D.

2025-12-08 bioinformatics
10.64898/2025.12.04.692308 bioRxiv
Show abstract

In clinical trials, it is common for only a subset of patients to respond to a given therapy. Such variability may arise from genetic differences, environmental influences, or the presence of distinct disease endotypes. Longitudinal transcriptomic profiles collected in phase 2a/2b studies provide a unique opportunity not only to investigate drug-induced biological mechanisms in humans, but also to understand why certain individuals fail to respond and to uncover previously unrecognized disease endotypes--ultimately informing the development of targeted therapeutics. However, analyzing these datasets is challenged by substantial patient heterogeneity, variability in disease severity at each visit, and the coarse temporal resolution due to sparse sampling. To address these limitations, we introduce a classification-guided autoencoder framework that jointly optimizes gene-expression reconstruction and classification objective to learn disease-relevant sample embeddings. Sample embeddings from all patients are leveraged to establish a continuous representation of disease dynamics, which can be clustered to delineate discrete disease states. We then construct a patient-sample graph in the learned latent space and apply a multi-commodity-flow based algorithm to infer patient trajectories through these states, enabling the identification of different patient trajectories. We evaluated our approach on three interventional datasets--ulcerative colitis, psoriasis, and atopic dermatitis, each containing both responders and non-responders. The method recapitulates known pathway perturbations associated with anti-IL17 and anti-IL6 therapies, identifies intermediate states reflecting disease/treatment progression, and reveals biologically meaningful patient trajectories within the patient population. Code and data are available in a public GitHub repository - https://github.com/Sanofi-Public/EndotypeDetection-Timeseries

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.