A Practice on Antibody Hydrophobic Interaction Chromatography Retention Time Prediction using Pre-Trained Large Language Model Fine-Tuning
Wang, B.; Cai, B.; Chen, H.; Xia, H.; Wang, B.; Liu, J.; Han, L.; Wang, R.
Show abstract
Hydrophobicity is a critical property associated with the risk of non-specific binding, and it is commonly assessed using hydrophobic interaction chromatography retention time. Several computational approaches have been developed to predict antibody developability based on pre-trained language models. Such models can be fine-tuned with limited labeled antibody sequences and, in principle, do not require structural information, which is often challenging to obtain. Nevertheless, few studies have achieved strong performance in hydrophobicity prediction without incorporating structural features. Here, we present a case study of fine-tuning the pre-trained model IgBert to predict antibody hydrophobicity. Using Herceptin as a reference, we performed hydrophobic interaction chromatography retention time experiments and generated Herceptin-adjusted datasets. The fine-tuned model achieved a best R2 of 0.916, underscoring the critical role of rigorous data quality control. We also synthesized and validated 20 commercially available antibody sequences, and the results showed that the predicted hydrophobic properties were correctly reflected. Our findings provide practical guidance and highlight considerations for future applications of fine-tuned pre-trained language models in antibody hydrophobicity prediction. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=189 HEIGHT=200 SRC="FIGDIR/small/742939v1_ufig1.gif" ALT="Figure 1"> View larger version (36K): org.highwire.dtl.DTLVardef@9c814eorg.highwire.dtl.DTLVardef@ed609dorg.highwire.dtl.DTLVardef@62172forg.highwire.dtl.DTLVardef@1e01d37_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Machine learning driven acceleration of biopharmaceutical formulation development using Excipient Prediction Software (ExPreSo) 91%
- DeepSCM: an efficient convolutional neural network surrogate model for the screening of therapeutic antibody viscosity 90%
- Insight on physicochemical properties governing peptide MS1 response in HPLC-ESI-MS/MS proteomics: A deep learning approach 90%
Similar papers in this journal
- Predicting purification process fit of monoclonal antibodies using machine learning 94%
- A high-throughput platform for biophysical antibody developability assessment to enable AI/ML model training 92%
- Enhancement of antibody thermostability and affinity by computational design in the absence of antigen 91%
Similar papers in this journal
- DeepDetect: deep learning of peptide detectability enhanced by peptide digestibility 92%
- Generalized calibration across LC-setups for generic prediction of small molecule retention times 92%
- Profiling Active Enzymes for Polysorbate Degradation in Biotherapeutics by Activity-Based Protein Profiling 91%
Similar papers in this journal
Similar papers in this journal
- Sequence-specific model for predicting peptide collision cross-section values in proteomic ion mobility spectrometry 92%
- Proteome Integral Stability Alteration assay dramatically increases throughput and sensitivity in profiling factor-induced proteome changes 92%
- PELSA-Decipher: a software tool for the processing and interpretation of ligand protein interaction dataset acquired by PELSA 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.