Back

AI-powered integration of multi-source data for TAA discovery to accelerate ADC and TCE drug development (I): TAA Target Identification and Prioritization

Xie, T.; Huang, C.-H.

2025-05-08 bioinformatics
10.1101/2025.05.06.652559 bioRxiv
Show abstract

The advancement of T-cell engagers (TCEs) and antibody-drug conjugates (ADCs) has been hindered by fragmented data landscapes. This paper, the first in a series, introduces an AI-driven framework specifically for tumor-associated antigen (TAA) target identification and prioritization, a critical initial step in TCE and ADC development. Our framework integrates diverse datasets-- including multi-omics repositories and information from scientific publications--to systematically enhance the discovery of TAAs. We have developed a graph retrieval-augmented generation (RAG)-enhanced language model that extracts insights from biological and clinical literature, while integrating curated public oncology-related omics databases such as TCGA, GTEx, single-cell atlases, and additional omics datasets. This approach prioritizes TAAs with high tumor selectivity and low on-target/off-tumor risk. By unifying diverse knowledge sources, our method provides a scalable, efficient, and data-agnostic strategy to address attrition challenges in both ADC and TCE drug development pipelines, focusing initially on TAA target identification and prioritization to transform the landscape of cancer therapeutics.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.