Back

Machine Learning for Predicting and Maximizing the Response of Breast Cancer Patients to Neoadjuvant Therapy

Chen, Z.; Cai, J. Y.

2025-10-14 oncology
10.1101/2025.10.11.25337587 medRxiv
Show abstract

PurposeNeoadjuvant therapy (NAT) is an established treatment for certain high-risk, locally advanced, or unresectable breast cancers, often facilitating breast-conserving surgery. Recent studies show that achieving pathologic complete response (pCR) after NAT correlates with higher event-free survival rates. Thus, accurate prediction of pCR is essential for personalizing breast cancer (BC) treatment to minimize side effects and improve effectiveness. MethodsWe used a machine learning model named XGBoost to predict pCR in BC patients. The classifier was trained on the expression values of 4,000 genes to predict pCR in ten arms of the I-SPY2 clinical trial. Based on these predictions, we developed a strategy to maximize pCR likelihood and identified influential genes using the importance scores from the model. ResultsXGBoost models for three arms, Pembrolizumab, ABT 888 plus carboplatin, and T-DM1 plus pertuzumab, achieved the highest prediction accuracies with areas under the receiver operation characteristic curve (AUCs) of 0.814, 0.792, and 0.788, respectively. If treatment assignments followed the XGBoost predictions, pCR rates for nine out of ten I-SPY2 arms could increase significantly, by 9.9% to 29.1% compared to trial results. Key genes associated with pCR were identified for each arm. The expression levels of some genes, including lower expression of ABDH1, AMZ1, BAIAP3 and SYTL4 and higher expression of DENND1C, HMGB3, HMMR, PLEKHF1 and RASEF, were associated with pCR in multiple arms. ConclusionThe machine learning models developed in this study provide accurate pCR predictions, improve pCR rates, and may find clinical applications to enhance the treatment for BC patients.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.