Back

Using Large Language Models to Assemble, Audit, and Prioritize the Therapeutic Landscape

Thyagatur Kidigannappa, A.; Sonti, N.; Vijayan, R.; Parikh, A.; Faller, R.

2025-11-16 bioinformatics
10.1101/2025.11.14.688562 bioRxiv
Show abstract

1We present an AI-assisted pipeline for disease-specific drug landscape analysis. Given a disease name, the system assembles a comprehensive, evidence-based view of therapeutic assets by integrating structured sources (such as ClinicalTrials.gov and ChEMBL) and unstructured sources (such as publications, press releases, and patents). Large language models are used in a constrained, auditable mode to normalize drug aliases, resolve drug-target/mechanism of action annotations, and harmonize program status across records. The output is a disease-centric map that spans preclinical assets, not-yet-approved assets (both active and discontinued/shelved), and FDA-approved drugs suitable for re-purposing. Assets are ranked using interpretable, evidence-based scoring heuristics that combine trial volume and clinical phase, endpoint outcomes, biomarker support, recency of activity, and regulatory designations, along with penalties for safety signals and non-pharmaceutical interventions, as well as proportional adjustments for operational versus scientific discontinuations. Case studies in Alzheimers disease, pancreatic cancer, and cystic fibrosis demonstrate generality, coverage, and discrimination across mechanisms and stages. This framework provides a transparent method to assemble and prioritize the therapeutic landscape for any disease, unifying disparate data into a coherent and analyzable representation.

Matching journals

The top 14 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.