Trace: A Fine-Tuned Biomedical Language Model For Directionally Informed Drug Repurposing From Transcriptome-Wide Association Studies
Otieno, C. O.; Seagle, H. M.; Akerele, A. T.; Jaworski, J.; Guare, L.; Setia-Verma, S.; Velez Edwards, D. R.; Edwards, T. L.
Show abstract
Transcriptome-wide association studies (TWAS) can identify genes where genetically predicted gene expression is associated with disease risk, but translating those signals into therapeutic opportunities remains time-consuming, manual, and difficult to reproduce. We developed TRACE (TWAS-driven Repurposing through AI-assisted Curation of Evidence), a gene- and phenotype-agnostic computational pipeline that accepts a TWAS gene and effect-size direction, normalizes the gene symbol, retrieves FDA-approved drug-gene candidates from four online resources, collects related peer-reviewed literature from PubMed, and uses a fine-tuned biomedical language model to classify whether the literature supports a direct drug-gene relationship, the mechanism of action, and the direction of effect. The pipeline then compares the drug-derived direction with the direction implied by the TWAS effect estimate to rank candidate therapeutic pairs and flag potential drug safety concerns. The local classifier, built on BiomedBERT, was trained using pipeline-derived labels, BioCreative VI ChemProt gold-standard chemical-protein relation examples, and author-reviewed active-learning cases, reaching a held-out macro F1 of 0.809 across three simultaneous classification tasks. We validated the pipeline against a manually curated endometriosis gold standard of 43 drug-gene pairs spanning six TWAS-identified genes, developed through S-PrediXcan analysis of endometriosis GWAS summary statistics, manual querying of four drug-gene interaction databases for each gene, literature review of drug-gene mechanistic evidence, and Mendelian randomization validation of candidate pairs. External validation used two independently published genetically informed drug-repurposing studies in metabolic dysfunction-associated steatotic liver disease (MASLD) and type 2 diabetes (T2D). The pipeline recovered 90.7% of endometriosis pairs, 88.2% of MASLD pairs, and 92.9% of T2D pairs that were present in at least one queried database. Applied to 99 endometriosis-associated TWAS genes, the pipeline identified 1,089 FDA-approved drug-gene pairs, 32 candidate therapeutic pairs, and 77 potential safety concerns, including independent recovery of leuprolide acetate, an established endometriosis therapy. This framework provides a scalable, literature-grounded bridge from TWAS discovery to prioritized therapeutic hypotheses, while preserving uncertainty through manual-review flags and requiring downstream Mendelian randomization, electronic health record-based validation, and experimental follow-up before clinical interpretation.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Multi-layered genetic approaches to identify approved drug targets 95%
- A practical guideline of genomics-driven drug discovery in the era of global biobank meta-analysis 93%
- Proteome-wide Mendelian randomization in global biobank meta-analysis reveals multi-ancestry drug targets for common diseases 92%
Similar papers in this journal
- scDrugPrio: A framework for the analysis of single-cell transcriptomics to address multiple problems in precision medicine in immune-mediated inflammatory diseases 92%
- Genome-wide prediction of pathogenic gain- and loss-of-function variants from ensemble learning of diverse feature set 91%
- From Text to Translation: Using Language Models to Prioritize Variants for Clinical Review 91%
Similar papers in this journal
Similar papers in this journal
- Community assessment of cancer drug combination screens identifies strategies for synergy prediction 94%
- Projecting genetic associations through gene expression patterns highlights disease etiology and drug mechanisms 93%
- In silico discovery of nanobody binders to a G-protein coupled receptor using AlphaFold-Multimer 92%
Similar papers in this journal
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 93%
- Few shot learning for phenotype-driven diagnosis of patients with rare genetic diseases 92%
- Federated Target Trial Emulation using Distributed Observational Data for Treatment Effect Estimation 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.