Hoike: A Joint-Embedding Predictive Architecture for Transcriptome Data Generation with Diffusion Models
Souza, P.; Ford, C. T.
Show abstract
In biomarker discovery, access to sufficient quantities of condition-specific transcriptomic data is often limited by cohort size, privacy concerns, and domain shift between normal and condition populations. Generative modeling can augment scarce cohorts and probe distributional transitions. Furthermore, synthetic transcriptome generation can support differential expression analyses, machine learning, privacy-preserving data sharing, benchmarking, and hypothesis generation in translational bioinformatics workloads in fields such as oncology. Here, we present Hoike, a framework that combines a crossdomain Joint-Embedding Predictive Architecture (JEPA) with a latent diffusion model to generate condition-specific bulk transcriptomes from a normal reference context. In Hoike, normal tissue profiles provide continuous conditioning signals, while the model learns disease-linked shifts in latent space and reconstructs gene-level expression in log2(TPM+1) space. The implementation supports paired normal-condition training, tissuealigned conditioning, and constrained non-negative decoding for biologically valid outputs. We describe the architecture, objective design, and evaluation protocol used in this work across GTEx-derived normal references and multiple TCGA condition cohorts as a case study. This serves as the technical specification of the Hoike framework and its reproducible analysis workflow.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics 94%
- Benchmarking computational methods for multi-omics biomarker discovery in cancer 94%
- PILOT-GM-VAE: Patient-Level Analysis of single cell Disease Atlas with Optimal Transport of Gaussian Mixture Variational Autoencoders 94%
Similar papers in this journal
- Identifying maximally informative signal-aware representations of single-cell data using the Information Bottleneck 94%
- An adversarial scheme for integrating multi-modal data on protein function 93%
- Learning multi-cellular representations of single-cell transcriptomics data enables characterization of patient-level disease states 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.