Back

Cross-Domain Transfer Learning from Peptides to Lipids Using a Multi-Property Fine-Tuned LLM

Anyaegbunam, U. A.; Teschner, D.; Schmidlin, T.; Hildebrandt, A.; Mayer, J. U.; Sprang, M.; Andrade, M.

2026-01-07 bioinformatics
10.64898/2026.01.06.697904 bioRxiv
Show abstract

Accurate knowledge of liquid chromatography retention time (RT) is essential for confident compound identification in metabolomics and lipidomics. Yet, it is often constrained by the scarcity of experimental data for many molecular classes. Current workflows depend on experimental RT libraries, which are time-consuming to build and limited to previously observed compounds. Here, we present a transfer learning pipeline that leverages large, publicly available peptide datasets to enable accurate lipid RT prediction in data-sparse scenarios. We first show that a ChemBERTa language model, when pre-trained on peptides with a multi-task objective (predicting both RT and fundamental RDKit molecular descriptors), learns a more robust and generalizable chemical representation than a single-task (RT-only) model. This multi-property pre-training yielded superior generalization in lipids, achieving test R2 values of 0.842 against 0.814 (RT-only). Crucially, transferring this peptide-based model to lipid data provided a pronounced advantage in data-sparse scenarios. When fine-tuned on only 5% of available lipid data, the transferred model improved the median test R2 by +0.234 over a model trained from scratch. Significant benefits persisted at intermediate data scales (50-75%), with performance converging only when 100% of the lipid data was used. Notably, the pre-trained model never underperformed the baseline, exhibiting more stable training across all data scales. These results demonstrate that multi-property pre-training guides language models towards chemically meaningful representations that support better RT prediction in different molecular domains. Furthermore, peptide-based pre-training facilitates cross-domain transfer of chemical properties to lipid. Our work provides a practical, scalable strategy to mitigate data scarcity in lipidomics by transferring knowledge from data-rich peptide databases, offering a computational alternative to extensive experimental library generation and enabling more confident identification in small-scale omics studies.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.