Back

Adversarial Validation Reveals Diagnostic Workflow Leakage in PCOS Machine Learning Models

Katarynczuk, K.; Stachowiak, A.; Piorkowska, N. J.; Ostromecki, A.; Franik, G.; Bizon, A.

2026-07-23 endocrinology
10.64898/2026.07.22.26358682 medRxiv
Show abstract

Background: Machine-learning models for polycystic ovary syndrome (PCOS) and other conditions frequently report near-perfect diagnostic performance, but retrospective datasets assembled from routine clinical practice can encode diagnostic-group membership in how data were acquired rather than in disease biology, and this acquisition-related information can be indistinguishable from genuine clinical signal under conventional validation. Objective: To determine, using a real-world PCOS cohort as a case study, whether high classification performance reflected clinically meaningful information or artifacts of data provenance, schema structure, and measurement-acquisition workflow, and to develop a generalizable audit framework for detecting such artifacts in retrospective medical machine learning. Methods: We analyzed 1,331 retrospective records (1,286 PCOS, 45 controls) from a single endocrine-gynecology database. A layered acquisition-bias framework compared classification performance using (i) raw and harmonized missingness patterns alone, (ii) measured values with and without explicit missingness indicators, and (iii) ascertainment-balanced feature sets with and without age. Logistic regression and random forest were evaluated using repeated stratified cross-validation, bootstrap resampling, label-permutation testing, and calibration analysis, and the framework was validated against a semi-synthetic experiment with known ground truth. Results: Diagnostic status was perfectly predicted (ROC-AUC = 1.000) from missingness patterns alone, before any clinical value was examined, and this persisted after semantic harmonization of duplicated source columns. Performance declined progressively as acquisition-sensitive information was removed, from near-ceiling in raw and harmonized value models to a mean ROC-AUC of approximately 0.80-0.82 in the most restrictive ascertainment-balanced, age-excluded representation. The semi-synthetic experiment reproduced this pattern under known data-generating conditions, confirming that harmonization removes schema-fragmentation artifacts but not workflow-driven acquisition bias. Conclusions: Apparent diagnostic performance in this cohort was substantially attributable to diagnostic workflow and data-acquisition structure rather than to a stable, transportable biological signal. The layered audit framework generalizes beyond PCOS and offers a practical tool for detecting acquisition-related leakage in retrospective clinical machine-learning studies.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.