AI-powered integration of multi-source data for TAA discovery to accelerate ADC and TCE drug development (I): TAA Target Identification and Prioritization
Xie, T.; Huang, C.-H.
Show abstract
The advancement of T-cell engagers (TCEs) and antibody-drug conjugates (ADCs) has been hindered by fragmented data landscapes. This paper, the first in a series, introduces an AI-driven framework specifically for tumor-associated antigen (TAA) target identification and prioritization, a critical initial step in TCE and ADC development. Our framework integrates diverse datasets-- including multi-omics repositories and information from scientific publications--to systematically enhance the discovery of TAAs. We have developed a graph retrieval-augmented generation (RAG)-enhanced language model that extracts insights from biological and clinical literature, while integrating curated public oncology-related omics databases such as TCGA, GTEx, single-cell atlases, and additional omics datasets. This approach prioritizes TAAs with high tumor selectivity and low on-target/off-tumor risk. By unifying diverse knowledge sources, our method provides a scalable, efficient, and data-agnostic strategy to address attrition challenges in both ADC and TCE drug development pipelines, focusing initially on TAA target identification and prioritization to transform the landscape of cancer therapeutics.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Knowledge Graph-based Thought: a knowledge graph enhanced LLMs framework for pan-cancer question answering 94%
- Strategies and Techniques for Quality Control and Semantic Enrichment with Multimodal Data: A Case Study in Colorectal Cancer with eHDPrep 94%
- MLcps: Machine Learning Cumulative Performance Score for classification problems 93%
Similar papers in this journal
- preon: Fast and accurate entity normalization for drug names and cancer types in precision oncology 95%
- AI-HOPE: An AI-Driven conversational agent for enhanced clinical and genomic data integration in precision medicine research 95%
- Variomes: a high recall search engine to support the curation of genomic variants 94%
Similar papers in this journal
Similar papers in this journal
- Mining drug-target interactions from biomedical literature using chemical and gene descriptions-based ensemble transformer model. 94%
- Identification of Monotonically Classifying Pairs of Genes for Ordinal Disease Outcomes 94%
- Prompt-to-Pill: Multi-Agent Drug Discovery and Clinical Simulation Pipeline 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.