Back

Phenotype-driven de novo molecular design from gene expression signatures

Xu, Y.; Kuang, T.; Ge, S.; Wu, H.; Wang, M.; Xu, H.; An, F.; Ma, Z.; Cheng, Q.; Ren, Z.

2026-07-24 bioinformatics
10.64898/2026.07.21.739736 bioRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWTarget-based and structure-guided drug design remain central to modern drug discovery, but complementary strategies are needed when predefined targets or binding pockets do not fully capture disease biology. Gene-expression signatures provide scalable system-level readouts of disease and perturbation states, making them attractive inputs for phenotype-guided molecular design. However, preserving phenotypic information during molecular generation remains challenging, and chemically plausible molecules may lose connection to the intended biological response. Here, we present Tx2Mol, a transcriptome-guided framework that translates gene-expression signatures into candidate molecules while maintaining biological guidance throughout generation. We evaluated Tx2Mol across three biological settings: bulk gene perturbation, single-cell perturbation, and patient-derived disease signatures; and three validation dimensions: chemical plausibility, structural compatibility, and phenotypic preservation. Across 10 cancer-relevant bulk gene-perturbation benchmarks, Tx2Mol outperformed 9 transcriptome-guided baselines, improving maximum Tanimoto similarity to known ligands by 24.10% on average and by 50.67% on HDAC1. Structure-based analyses further supported structurally novel candidates with favorable predicted target binding. Tx2Mol also generalized to noisy single-cell perturbation profiles and preserved drug-induced transcriptional responses through in silico drug-perturbation validation. Patient-derived disease signatures further guided molecular generation toward approved-drug chemical space. Together, these results support gene-expression phenotypes as actionable guidance signals for phenotype-directed molecular design and candidate prioritization.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.