Using Large Language Models to Assemble, Audit, and Prioritize the Therapeutic Landscape
Thyagatur Kidigannappa, A.; Sonti, N.; Vijayan, R.; Parikh, A.; Faller, R.
Show abstract
1We present an AI-assisted pipeline for disease-specific drug landscape analysis. Given a disease name, the system assembles a comprehensive, evidence-based view of therapeutic assets by integrating structured sources (such as ClinicalTrials.gov and ChEMBL) and unstructured sources (such as publications, press releases, and patents). Large language models are used in a constrained, auditable mode to normalize drug aliases, resolve drug-target/mechanism of action annotations, and harmonize program status across records. The output is a disease-centric map that spans preclinical assets, not-yet-approved assets (both active and discontinued/shelved), and FDA-approved drugs suitable for re-purposing. Assets are ranked using interpretable, evidence-based scoring heuristics that combine trial volume and clinical phase, endpoint outcomes, biomarker support, recency of activity, and regulatory designations, along with penalties for safety signals and non-pharmaceutical interventions, as well as proportional adjustments for operational versus scientific discontinuations. Case studies in Alzheimers disease, pancreatic cancer, and cystic fibrosis demonstrate generality, coverage, and discrimination across mechanisms and stages. This framework provides a transparent method to assemble and prioritize the therapeutic landscape for any disease, unifying disparate data into a coherent and analyzable representation.
Matching journals
The top 14 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Scoping Review of Artificial Intelligence Applications in Clinical Trial Risk Assessment 94%
- Federated Target Trial Emulation using Distributed Observational Data for Treatment Effect Estimation 93%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 93%
Similar papers in this journal
- Analysis of clinical trial registry entry histories using the novel R package cthist 93%
- Combining explainable machine learning, demographic and multi-omic data to identify precision medicine strategies for inflammatory bowel disease 93%
- TargetDB: A target information aggregation tool and tractability predictor 92%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Data-driven strategies for drug repurposing: insights, recommendations, and case studies 94%
- OncoPubMiner: A platform for oncology publication mining 91%
- PheCode-guided multi-modal topic modeling of electronic health records improves disease incidence prediction and GWAS discovery from UK Biobank 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.