Back

Overcoming Topology Bias and Cold-Start Limita-tions in Drug Repurposing: A Clinical-Outcome-Aligned LLM Framework

DU, R.; FUNG, M.; HU, Y.; LIU, D.

2026-01-13 pharmacology and toxicology
10.64898/2026.01.12.699148 bioRxiv
Show abstract

Graph Neural Networks (GNNs) in drug repurposing suffer from two limitations: transductive failure in zero-shot (cold-start) scenarios and popularity bias that misidentifies high-degree nodes as effective drugs. We propose a framework that shifts the optimization objective from graph topology to clinical utility, integrating Knowledge Graph RAG (KG-RAG), Supervised Fine-Tuning (SFT), and Kahneman-Tversky Optimization (KTO) using Phase III clinical trial outcomes as rewards. We evaluated our approach on a rigorous 1:10 negative sampling benchmark derived from MiRAGE, covering Standard, Cold-Start, and Degree-Matched settings. In Cold-Start scenarios where topological signals are absent, traditional GNNs (including TxGNN) collapse (Top-10 Precision < 0.30), whereas DR-SFT model achieves 0.80, demonstrating robust semantic generalization for novel compounds. Crucially, in Degree-Matched tests dominated by "popular but ineffective" decoys, the DR-KTO acts as a clinical gatekeeper, achieving 0.90 Top-10 Precision--significantly outperforming DR-SFT (0.70) and GNNs (0.2-0.4) by effectively penalizing hard negatives. Beyond repurposing accuracy, the model achieves state-of-the-art performance on BioASQ, and increased ability in Chemprot. Orthogonal validation via DrugReAlign confirms physical plausibility, yielding significantly lower docking binding energies for KTO-recommended candidates. SPR experiments further corroborate these findings. By aligning LLM reasoning with clinical evidence, our framework successfully bridges the gap between semantic inference, topological structure, and clinical reality.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.