Back

From bench assays to bedside: context-embedding transformer predicts monoclonal antibody viscosity, clearance, and regulatory success

Virk, S. S.; Virk, A. S.

2025-11-01 bioinformatics
10.1101/2025.10.31.685722 bioRxiv
Show abstract

Formulation and pharmacokinetic liabilities remain major bottlenecks in monoclonal antibody development. Here, we present the agnostic context-embedding transformer (ACeT), an interpretable machine-learning framework that integrates heterogeneous early assay panels into endpoint-specific developability models. Using published antibody datasets, ACeT predicted high-concentration viscosity from four dilute-solution assays with held-out R{superscript 2} {approx} 0.75 and root-mean-square error {approx} 4.8 cP across a 5-45 cP range. From four clearance-related in vitro assays, it predicted mouse intravenous exposure with held-out R{superscript 2} = 0.80 and normalized root-mean-square error = 0.15. On a public 152-antibody panel, ACeT predicted hydrophobic interaction chromatography retention time, an orthogonal stickiness/hydrophobicity readout, with out-of-fold Pearson r{superscript 2} {approx} 0.83 and outperformed a published quantitative structure-property relationship baseline. In an exploratory retrospective analysis using five early developability assays, ACeT classified clinical outcomes (Approved vs Terminated) with balanced accuracy of [~]0.78 on a held-out internal set of 23 clinical IgG1 antibodies with outcome-locked labels and 0.83 on a temporally independent external cohort of 14 antibodies. Feature attribution recovered mechanistically plausible drivers, including the diffusion interaction parameter kD and size-exclusion chromatography peak-shape metrics for viscosity and heparin-, baculovirus particle-, and poly-D-lysine- signals for exposure. These results show that routine assay panels can support practical, interpretable machine-learning guided triage of antibodies for formulation and pharmacokinetic risk, and may capture a developability-linked component of downstream progression risk.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.