Identification of Patient Trajectories in Timeseries Clinical Transcriptomics Data
Mathur, S.; Chen, Y.; Passban, P.; Zhang, L.; Habiel, D.
Show abstract
In clinical trials, it is common for only a subset of patients to respond to a given therapy. Such variability may arise from genetic differences, environmental influences, or the presence of distinct disease endotypes. Longitudinal transcriptomic profiles collected in phase 2a/2b studies provide a unique opportunity not only to investigate drug-induced biological mechanisms in humans, but also to understand why certain individuals fail to respond and to uncover previously unrecognized disease endotypes--ultimately informing the development of targeted therapeutics. However, analyzing these datasets is challenged by substantial patient heterogeneity, variability in disease severity at each visit, and the coarse temporal resolution due to sparse sampling. To address these limitations, we introduce a classification-guided autoencoder framework that jointly optimizes gene-expression reconstruction and classification objective to learn disease-relevant sample embeddings. Sample embeddings from all patients are leveraged to establish a continuous representation of disease dynamics, which can be clustered to delineate discrete disease states. We then construct a patient-sample graph in the learned latent space and apply a multi-commodity-flow based algorithm to infer patient trajectories through these states, enabling the identification of different patient trajectories. We evaluated our approach on three interventional datasets--ulcerative colitis, psoriasis, and atopic dermatitis, each containing both responders and non-responders. The method recapitulates known pathway perturbations associated with anti-IL17 and anti-IL6 therapies, identifies intermediate states reflecting disease/treatment progression, and reveals biologically meaningful patient trajectories within the patient population. Code and data are available in a public GitHub repository - https://github.com/Sanofi-Public/EndotypeDetection-Timeseries
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deep feature extraction of single-cell transcriptomes by generative adversarial network 94%
- High-dimensional Biomarker Identification for Scalable and Interpretable Disease Prediction via Machine Learning Models 94%
- Looking at the BiG picture: Incorporating bipartite graphs in drug response prediction 94%
Similar papers in this journal
- Novel multi-omics deconfounding variational autoencoders can obtain meaningful disease subtyping 94%
- PheCode-guided multi-modal topic modeling of electronic health records improves disease incidence prediction and GWAS discovery from UK Biobank 94%
- GexMolGen: Cross-modal Generation of Hit-like Molecules via Large Language Model Encoding of Gene Expression Signatures 94%
Similar papers in this journal
- A Generalized Higher-order Correlation Analysis Framework for Multi-Omics Network Inference 94%
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 94%
- A variational autoencoder trained with priors from canonical pathways increases the interpretability of transcriptome data 94%
Similar papers in this journal
- SwarmMAP: Swarm Learning for Decentralized Cell Type Annotation in Single Cell Sequencing Data 94%
- Spatial cell graph analysis reveals skin tissue organization characteristic for cutaneous T cell lymphoma 93%
- A personalised approach for identifying disease-relevant pathways in heterogeneous diseases 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.