Early prediction of ovarian cancer risk based on real world data
de la Oliva, V.; Esteban-Medina, A.; Alejos, L.; Munoyerro-Muniz, D.; Villegas, R.; Dopazo, J.; Loucera, C.
Show abstract
This study presents the development of an early prediction model for high-grade serous ovarian cancer (HGSOC) using real-world data from the Andalusian Health Population Database (BPS), containing electronic health records (EHR) of over 15 million patients. Leveraging the extensive data availability, the model aims to identify individuals at high risk of HGSOC without the need for specific tumor markers or prior stratification into risk groups. Utilizing an Explainable Boosting Machine (EBM) algorithm, the model incorporates diverse clinical variables including demographics, chronic diseases, symptoms, blood test results, and healthcare utilization patterns. The model was trained and validated using a total of 3,088 HGSOC patients diagnosed between 2018 and 2022 along with 114,942 controls of similar characteristics, to emulate the prevalence of the disease, achieving a sensitivity of 0.65 and a specificity of 0.85. This study underscores the importance of using patient data from the general population, demonstrating that effective early detection models can be developed from routinely collected healthcare data. The approach addresses limitations of traditional screening methods by providing a cost-effective and broadly applicable tool for early cancer detection, potentially improving patient outcomes through timely interventions. The interpretability of the early prediction model also offers insights into the most significant predictors of cancer risk, further enhancing its utility in clinical settings.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Bayesian combination of mechanistic modeling and machine learning (BaM3): improving personalized tumor growth predictions 91%
- Machine learning models predict long COVID outcomes based on baseline clinical and immunologic factors 91%
- Subpopulation-specific Machine Learning Prognosis for Underrepresented Patients with Double Prioritized Bias Correction 91%
Similar papers in this journal
- Subtyping of common complex diseases and disorders by integrating heterogeneous data. Identifying clusters among women with lower urinary tract symptoms in the LURN study 94%
- Identification of high-risk COVID-19 patients using machine learning 94%
- ChatGPT-Enhanced ROC Analysis (CERA): A Shiny Web Tool for Finding Optimal Cutoff in Biomarker Analysis 93%
Similar papers in this journal
- Mitigating Machine Learning Bias Between High Income and Low-Middle Income Countries for Enhanced Model Fairness and Generalizability 94%
- On evaluation metrics for medical applications of artificial intelligence 93%
- Novel ratio-metric features enable the identification of new driver genes across cancer types 93%
Similar papers in this journal
Similar papers in this journal
- A regularized functional regression model enabling transcriptome-wide dosage-dependent association study of cancer drug response 93%
- Revealing cancer driver genes through integrative transcriptomic and epigenomic analyses with Moonlight 92%
- Using random forests to uncover the predictive power of distance-varying cell interactions in tumor microenvironments 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.