Adversarial Validation Reveals Diagnostic Workflow Leakage in PCOS Machine Learning Models
Katarynczuk, K.; Stachowiak, A.; Piorkowska, N. J.; Ostromecki, A.; Franik, G.; Bizon, A.
Show abstract
Background: Machine-learning models for polycystic ovary syndrome (PCOS) and other conditions frequently report near-perfect diagnostic performance, but retrospective datasets assembled from routine clinical practice can encode diagnostic-group membership in how data were acquired rather than in disease biology, and this acquisition-related information can be indistinguishable from genuine clinical signal under conventional validation. Objective: To determine, using a real-world PCOS cohort as a case study, whether high classification performance reflected clinically meaningful information or artifacts of data provenance, schema structure, and measurement-acquisition workflow, and to develop a generalizable audit framework for detecting such artifacts in retrospective medical machine learning. Methods: We analyzed 1,331 retrospective records (1,286 PCOS, 45 controls) from a single endocrine-gynecology database. A layered acquisition-bias framework compared classification performance using (i) raw and harmonized missingness patterns alone, (ii) measured values with and without explicit missingness indicators, and (iii) ascertainment-balanced feature sets with and without age. Logistic regression and random forest were evaluated using repeated stratified cross-validation, bootstrap resampling, label-permutation testing, and calibration analysis, and the framework was validated against a semi-synthetic experiment with known ground truth. Results: Diagnostic status was perfectly predicted (ROC-AUC = 1.000) from missingness patterns alone, before any clinical value was examined, and this persisted after semantic harmonization of duplicated source columns. Performance declined progressively as acquisition-sensitive information was removed, from near-ceiling in raw and harmonized value models to a mean ROC-AUC of approximately 0.80-0.82 in the most restrictive ascertainment-balanced, age-excluded representation. The semi-synthetic experiment reproduced this pattern under known data-generating conditions, confirming that harmonization removes schema-fragmentation artifacts but not workflow-driven acquisition bias. Conclusions: Apparent diagnostic performance in this cohort was substantially attributable to diagnostic workflow and data-acquisition structure rather than to a stable, transportable biological signal. The layered audit framework generalizes beyond PCOS and offers a practical tool for detecting acquisition-related leakage in retrospective clinical machine-learning studies.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Improving Pre-eclampsia Risk Prediction by Modeling Individualized Pregnancy Trajectories Derived from Routinely Collected Electronic Medical Record Data 91%
- Zero-shot Interpretable Phenotyping of Postpartum Hemorrhage Using Large Language Models 91%
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 91%
Similar papers in this journal
- Urinary tract infections in children: building a causal model-based decision support tool for diagnosis with domain knowledge and prospective data 90%
- Quantitative bias analysis in practice: Review of software for regression with unmeasured confounding 90%
- Towards reduction in bias in epidemic curves due to outcome misclassification through Bayesian analysis of time-series of laboratory test results: Case study of COVID-19 in Alberta, Canada and Philadelphia, USA 89%
Similar papers in this journal
- Deep plasma proteomics identifies and validates an eight-protein biomarker panel that separate benign from malignant tumors in ovarian cancer 90%
- Estimating Heritability of Glycaemic Response to Metformin using Nationwide Electronic Health Records and Population-Sized Pedigree 90%
- A user-friendly tool for cloud-based whole slide image segmentation, with examples from renal histopathology 90%
Similar papers in this journal
Similar papers in this journal
- Who is pregnant? defining real-world data-based pregnancy episodes in the National COVID Cohort Collaborative (N3C) 92%
- Clinical interpretation of machine learning models for prediction of diabetic complications using electronic health records 90%
- Modeling physician variability to prioritize relevant medical record information 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.