A Cross-Cohort Validated Plasma Lipid Biomarker Assay for Early Breast Cancer Detection Using Machine Learning
Huang, T.; Koch, F. C.; Peake, D. A.; Adam, K.-P.; David, M.; Li, D.; Heffernan, K.; Lim, A.; Hurrell, J. G.; Preston, S.; Baterseh, A.; Vafaee, F.
Show abstract
Early detection of breast cancer remains essential for improving clinical outcomes, and complementary non-invasive approaches are needed to support existing screening methods, particularly for women with dense breast tissue. We have previously reported plasma lipid biomarker discovery using untargeted high-resolution liquid chromatography tandem mass spectrometry (LC-MS/MS). In this study, we performed biomarker confirmation and developed machine-learning models applied to targeted plasma lipid measurements for the non-invasive detection of early-stage breast cancer across international cohorts with independent external validation. Targeted LC-MS/MS was used to quantify candidate lipid panels in plasma samples from European discovery cohorts (n = 554) and an independent Australian cohort (n = 266) used for external validation. Data-driven feature selection identified a 15-lipid panel with strong performance in European cohorts (AUC [≥] 0.94). External validation prior to confidence stratification yielded 76% sensitivity, 64% specificity, and an AUC of 0.81 in the Australian validation cohort. Clinical assay development requires iterative panel and model testing to support translational feasibility and performance in the intended-use population. An analytically viable panel, excluding lipids requiring complex and costly synthesis, achieved comparable accuracy with improved assay robustness. Confidence-based analysis showed enhanced performance for predictions made with moderate to high confidence, with sensitivity up to 89% and AUC up to 0.85, suggesting that ongoing research should focus on strategies to enhance diagnostic model confidence. Importantly, model predictions were independent of breast density, tumour size, grade, subtype, and morphology, indicating biological specificity of the lipid signature. These results demonstrate that calibrated machine-learning models applied to plasma lipid biomarkers can support non-invasive breast cancer detection. Expanding training datasets to include greater diversity will further improve performance in the ongoing development of this lipid-based detection approach.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A new pipeline for the normalization and pooling of metabolomics data 92%
- Screening for inborn errors of metabolism using untargeted metabolomics and out-of-batch controls 92%
- Do mass-spectrometry-derived metabolomics improve prediction of pregnancy-related disorders? Findings from a UK birth cohort with independent validation 92%
Similar papers in this journal
Similar papers in this journal
- Deep plasma proteomics identifies and validates an eight-protein biomarker panel that separate benign from malignant tumors in ovarian cancer 92%
- Pervasive Influence of Hormonal Contraceptives on the Human Plasma Proteome in a Broad Population Study 91%
- Subpopulation-specific Machine Learning Prognosis for Underrepresented Patients with Double Prioritized Bias Correction 90%
Similar papers in this journal
- The significance of molecular heterogeneity in breast cancer batch correction and dataset integration 95%
- Automated quantification of Ki-67 expression in breast cancer from H&E-stained slides using a transformer-based regression model 92%
- Multiomic profiling of metastatic potential in estrogen receptor-positive human epidermal growth factor-negative breast cancer 92%
Similar papers in this journal
- Analytical Validation of MyProstateScore 2.0 90%
- A Machine Learning Ensemble Based on Radiomics to Predict BI-RADS Category and Reduce the Biopsy Rate of Ultrasound-Detected Suspicious Breast Masses 90%
- A clinical MALDI-ToF Mass spectrometry assay for SARS-CoV-2: Rational design and multi-disciplinary team work. 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.