Back

CancerSpot: A multi-cancer early detection test developed and validated on a retrospective cohort

Basu, S.; Hiremath, P.; Rathod, N.; Chatterjee, A.; Vishwanath, D.; Ghosh, A.; Sanguri, S.; Chakraborty, S.; Tripathi, A.; RT, P.; Nair, A.; Kumar, G.; Sekar, K.; Yete, S.; G, B.; Bahadur, U.; Radhakrishnan, A.; Khan, A.; Kannan S, Y.; Bollipalli, L.; Ghana, P.; Ramanathan, A.; Saha, P.; Phalke, S.; Cantor, C.; Limaye, S.; Chandru, V.; Veeramachaneni, V.; Hariharan, R.

2024-12-05 oncology
10.1101/2024.12.03.24318395 medRxiv
Show abstract

Next-generation sequencing (NGS) technologies have transformed biomarker discovery, enabling the detection of disease-associated markers at the earliest stages of illness. In this study, we introduce a blood-based, non-invasive test for multi-cancer detection using cell-free DNA (cfDNA) methylation sequencing. The test employs a novel methylation scoring system derived from sequencing data and integrates machine learning to analyze a retrospective cohort of newly diagnosed cancer cases and controls recruited from multiple centers across India. To enhance robustness, the study includes a substantial proportion of controls with habitual tobacco and alcohol use, ensuring the tests resilience against confounding factors. The tests accuracy was further validated through synthetic data augmentation, demonstrating reliability under conditions of random signal perturbation. At an approximate specificity of 97%, the assay achieves sensitivities of 79.3% for Stage I, 78.4% for Stage II, 78.4% for Stage III, and 86.8% for Stage IV cancers in an independent validation cohort. Additionally, the test demonstrates Top 2 Tissue of Origin (TOO) accuracies of 78.3% for Stage I, 79.3% for Stage II, 82.8% for Stage III, and 69.7% for Stage IV cancers. This blood-based test holds considerable promise for early cancer detection, offering a precise test for cancer screening.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.