Hierarchical Representation Learning for Drug Mechanism-of-Action Prediction from Gene Expression Data
Katsaouni, N.; Schulz, M. H.
Show abstract
Deciphering drug mechanisms of action (MoAs) from transcriptional responses is key for discovery and repurposing. While recent machine learning approaches improve prediction accuracy beyond traditional similarity metrics, they often lack biological structure and interpretability in the learned space. We introduce a hierarchical representation learning framework that explicitly enforces mechanistically coherent organization using dual ArcFace objectives, yielding an interpretable latent space that captures both MoA-level separation and compound-level substructure. Gene importance and pathway enrichment analyses confirm that the learned representations recover established signaling programs. Trained on LINCS L1000 data, the model also improves F1 performance over state-of-the-art baselines and generalizes to unseen compounds and cell types. Additionally, the latent space generalizes to CRISPR knockdowns without the need for retraining, indicating it captures pathway-level perturbations independently of modality.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predicting drug polypharmacology from cell morphology readouts using variational autoencoder latent space arithmetic 97%
- Capturing cell heterogeneity in representations of cell populations for image-based profiling using contrastive learning 95%
- Impact of between-tissue differences on pan-cancer predictions of drug sensitivity 94%
Similar papers in this journal
- First fully-automated AI/ML virtual screening cascade implemented at a drug discovery centre in Africa 95%
- comboFM: leveraging multi-way interactions for systematic prediction of drug combination effects 95%
- Community assessment of cancer drug combination screens identifies strategies for synergy prediction 94%
Similar papers in this journal
- TrustAffinity: accurate, reliable and scalable out-of-distribution protein-ligand binding affinity prediction using trustworthy deep learning 95%
- Sagittarius: Extrapolating Heterogeneous Time-Series Gene Expression Data 94%
- An interpretable deep learning framework for genome-informed precision oncology 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.