Transcriptomic data and biomedical literature synergize in finding pharmacologic gene regulators
Deisseroth, C. A.; Brazelton, B.; Shaik, Z.; Liu, Z.; Zoghbi, H. Y.
Show abstract
Most Mendelian disorders caused by a deficiency or excess of one gene product lack targeted therapies. Since these disorders can be modeled with a gene overexpression, knockout, or knockdown, drugs that oppose the transcriptomic effects of such perturbations may be promising therapeutic candidates. RNA-Sequencing (RNA-Seq) studies can fuel this drug-prioritization, but their labels, written in plain language, must be annotated manually. Hence, we introduce Signature-based Networks from Automatically Curated Knockout, Knockdown, and Small-molecule Studies (SNACKKSS), which automatically curates gene-disruption and drug studies from the Gene Expression Omnibus and, in partnership with uniformly computed read count datasets, feeds the labels and RNA-Seq data directly into regulatory relationship predictions. Through cross-validation, we show that SNACKKSS predictions (specifically, from a variation called "SA4") make a unique contribution to finding protein-inhibiting compounds, even alongside existing predictors. We demonstrate the benefit of aggregating multiple predictive tools, and provide this powerful ensemble alongside SNACKKSS. Importantly, we advise researchers to test complex machine learning models on multiple devices. Even with code packages kept consistent, they can run deterministically within a machine, but inconsistently on different ones. Nonetheless, the downstream predictive ability was striking, and leveraging multiple sources of information, RNA-Seq data included, will vastly improve drug-repurposing screens.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Specifying cellular context of transcription factor regulons for exploring context-specific gene regulation programs 95%
- Differential Expression Enrichment Tool (DEET): an interactive atlas of human differential gene expression 95%
- Supervised learning with word embeddings derived from PubMed captures latent knowledge about protein kinases and cancer 94%
Similar papers in this journal
- GRAND: A database of gene regulatory network models across human conditions 97%
- Signatures of cell death and proliferation in perturbation transcriptomics data - from confounding factor to effective prediction 95%
- noisyR: Enhancing biological signal in sequencing datasets by characterising random technical noise 95%
Similar papers in this journal
- GeneWalk identifies relevant gene functions for a biological context using network representation learning 95%
- Robust differential expression testing for single-cell CRISPR screens at low multiplicity of infection 95%
- Partial gene suppression improves identification of cancer vulnerabilities when CRISPR-Cas9 knockout is pan-lethal 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.