Cross-Domain Transfer Learning from Peptides to Lipids Using a Multi-Property Fine-Tuned LLM
Anyaegbunam, U. A.; Teschner, D.; Schmidlin, T.; Hildebrandt, A.; Mayer, J. U.; Sprang, M.; Andrade, M.
Show abstract
Accurate knowledge of liquid chromatography retention time (RT) is essential for confident compound identification in metabolomics and lipidomics. Yet, it is often constrained by the scarcity of experimental data for many molecular classes. Current workflows depend on experimental RT libraries, which are time-consuming to build and limited to previously observed compounds. Here, we present a transfer learning pipeline that leverages large, publicly available peptide datasets to enable accurate lipid RT prediction in data-sparse scenarios. We first show that a ChemBERTa language model, when pre-trained on peptides with a multi-task objective (predicting both RT and fundamental RDKit molecular descriptors), learns a more robust and generalizable chemical representation than a single-task (RT-only) model. This multi-property pre-training yielded superior generalization in lipids, achieving test R2 values of 0.842 against 0.814 (RT-only). Crucially, transferring this peptide-based model to lipid data provided a pronounced advantage in data-sparse scenarios. When fine-tuned on only 5% of available lipid data, the transferred model improved the median test R2 by +0.234 over a model trained from scratch. Significant benefits persisted at intermediate data scales (50-75%), with performance converging only when 100% of the lipid data was used. Notably, the pre-trained model never underperformed the baseline, exhibiting more stable training across all data scales. These results demonstrate that multi-property pre-training guides language models towards chemically meaningful representations that support better RT prediction in different molecular domains. Furthermore, peptide-based pre-training facilitates cross-domain transfer of chemical properties to lipid. Our work provides a practical, scalable strategy to mitigate data scarcity in lipidomics by transferring knowledge from data-rich peptide databases, offering a computational alternative to extensive experimental library generation and enabling more confident identification in small-scale omics studies.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Beyond the Leaderboard: Leveraging Predictive Modeling for Protein-Ligand Insights and Discovery 94%
- CaLMPhosKAN: Prediction of General Phosphorylation Sites in Proteins via Fusion of Codon Aware Embeddings with Amino Acid Aware Embeddings and Wavelet-based Kolmogorov Arnold Network 94%
- BERTMHC: Improves MHC-peptide class II interaction prediction with transformer and multiple instance learning 94%
Similar papers in this journal
Similar papers in this journal
- PLMFit : Benchmarking Transfer Learning with Protein Language Models for Protein Engineering 95%
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 94%
- AI-Guided Discovery and Optimization of Antimicrobial Peptides Through Species-Aware Language Model 93%
Similar papers in this journal
- Joint structural annotation of small molecules using liquid chromatography retention order and tandem mass spectrometry data 95%
- Annotating metabolite mass spectra with domain-inspired chemical formula transformers 94%
- Evaluating generalizability of artificial intelligence models for molecular datasets 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.