Back

Diagnostics

MDPI AG

Preprints posted in the last 30 days, ranked by how well they match Diagnostics's content profile, based on 50 papers previously published here. The average preprint has a 0.08% match score for this journal, so anything above that is already an above-average fit.

1
Machine learning models for predicting prostate cancer and clinically significant prostate cancer at biopsy: An updated analysis of an expanded Japanese cohort

Takeuchi, T.; Nomiya, A.

2026-08-11 urology 10.64898/2026.08.09.26360024 medRxiv
Top 0.1%
23.3%
Show abstract

Background: A 2019 report from our institution described a multilayer artificial neural network (ANN) for predicting prostate cancer at biopsy in 334 patients, trained with TensorFlow 1.x and evaluated at three fixed step counts without separating hyperparameter selection from test evaluation. We re-analyzed an expanded cohort from the same institution using contemporary machine-learning practice. Methods: We pooled all available biopsy episodes from the same institutional database (n = 526; 524 after excluding one non-binary outcome code and one record with missing digital rectal examination [DRE] data), retaining the same seven predictors used in the original report (age, prior biopsy history, PSA, prostate volume, DRE, and MRI diffusion-weighted imaging findings in the peripheral and transition zones). Because 27 patients contributed more than one biopsy episode, we used patient-ID-grouped, stratified k-fold cross-validation (StratifiedGroupKFold; scikit-learn 1.8.0) with 3 and 5 folds, repeated over 10 random partitions, to avoid leakage between folds. Four classifiers were compared: L2-regularized logistic regression, gradient boosting, random forest, and a shallow (single hidden layer) multilayer perceptron. Two outcomes were modeled: detection of any prostate cancer, and detection of clinically significant prostate cancer (Gleason score [≥] 7). Results: Any-cancer prevalence was 55.7% (292/524) and Gleason score [≥] 7 prevalence was 39.7% (208/524). With repeated 5-fold cross-validation, gradient boosting gave the highest discrimination for any prostate cancer (mean AUC 0.826, 95% CI 0.823-0.830) and for Gleason score [≥] 7 (mean AUC 0.855, 95% CI 0.852-0.859), closely followed by random forest and logistic regression (AUC 0.81-0.85). The shallow multilayer perceptron performed worse and less consistently than the other three models (any-cancer AUC 0.671; Gleason score [≥] 7 AUC 0.742) and than the deeper five-hidden-layer ANN reported in 2019. Results with 3-fold cross-validation were essentially unchanged. Conclusions: In an expanded cohort, regularized logistic regression, gradient boosting, and random forest all discriminated prostate cancer at biopsy at least as well as the previously reported multilayer ANN, using far simpler models and a methodology that separates hyperparameter tuning from performance estimation. A shallow neural network offered no advantage over these simpler alternatives in this sample size. This is a preprint; the study has not undergone external peer review.

2
Modeling Biomarker-Guided Avoidance of Radical Cystectomy: Costs and Outcomes

Sholklapper, T. N.; Li, M.; Srivastava, A.; Wagh, A.; Handorf, E.; Beck, J. R.; Abbosh, P.

2026-08-19 urology 10.64898/2026.08.17.26360582 medRxiv
Top 0.1%
11.5%
Show abstract

Importance There is growing interest to avoid radical cystectomy (RC) in patients with muscle-invasive bladder cancer (MIBC) who receive neoadjuvant chemotherapy and achieve pathological complete response (ypCR). To achieve this goal, molecular biomarkers will likely need to be used to enhance clinical staging given the limitations of evaluation by cystoscopy, cytology, and cross-sectional imaging. There are no studies evaluating whether safe RC avoidance (SaRCA) would be cost effective and what impacts it would have on quality of life (QoL) and survival. Objective This study models the potential economic, QoL, and survival costs/benefits of a ypCR biomarker as it relates to SaRCA using a decision analysis and Markov Model (MM). Methods/Materials A decision tree and MM was created to compare the expected costs of initial treatment, QoL, and survival under one strategy where all patients undergo RC after neoadjuvant treatment versus an alternative strategy where all patients would be subjected to the biomarker test with biomarker-positive patients (those with presumed residual disease) undergoing RC, while biomarker-negative patients (presumed complete responders) would undergo surveillance for up to 20 years. ypCR rates to neoadjuvant therapy, survival with and without RC, quality adjusted life years (QALY), and costs were abstracted from the literature. Test cost, sensitivity, and specificity were also abstracted from the literature for multiple clinical or liquid biopsy approaches. Results Broadly, SaRCA approaches are cost effective with the exception of systematic endoscopic evaluation (SEE). All testing approaches result in higher QALY and overall life expectancy compared to no testing. The cost of the test is offset by decreased usage of RC to realize a cost savings. These domains are further improved when cisplatin-based chemotherapy is replaced with emerging neoadjuvant therapies. Conclusions and relevance Modeling supports the development of accurate biomarker tests which can distinguish residual disease states to enable SaRCA. Such a biomarker could be used to avoid an expensive and risky operation, and unexpectedly would provide a survival benefit by reducing the number of perioperative mortalities in patients achieving ypCR. Development of an accurate biomarker-based test is likely to reduce cost and increase QoL and survival. An accurate biomarker test would have utility for patients, payers, hospitals, and physicians.

3
Multimodal Radiogenomic Machine Learning for Biochemical Recurrence Prediction Following Radical Prostatectomy Using PSMA-PET, mpMRI, and the Decipher Genomic Classifier

Reddy Chimmula, R.; Yong, C.; Love, H. L.; Shiradkar, R.; Holmes, J.; Nair, V.; Tann, M.; Bahler, C.; Oderinde, O. M.

2026-08-10 urology 10.64898/2026.08.05.26359806 medRxiv
Top 0.1%
10.4%
Show abstract

Background: Biochemical recurrence (BCR) occurs in up to 40% of men following radical prostatectomy (RP). Current risk models rely primarily on clinicopathologic variables and may not fully capture the biological heterogeneity associated with recurrence. The Decipher Genomic Classifier (DGC), prostate-specific membrane antigen positron emission tomography (PSMA-PET), and multiparametric magnetic resonance imaging (mpMRI) provide complementary prognostic information that may improve prediction. Objective: To develop and evaluate machine learning (ML) models integrating DGC, PSMA-PET, and mpMRI for preoperative prediction of BCR following RP. Methods: This retrospective study included patients with available preoperative DGC, PSMA-PET, mpMRI, and clinicopathologic data. Logistic regression (LR), random forest (RF), and XGBoost models were developed using single- and multimodality feature combinations. Early- and intermediate-fusion strategies were evaluated. Performance was assessed using an area under the receiver operating characteristic curve (AUC) and accuracy. Clinical utility was evaluated using decision curve analysis. Results: XGBoost consistently outperformed LR and RF. DGC achieved the highest single-modality performance (AUC 0.94, accuracy 86.7%). Among multimodal models, DGC combined with PSMA-PET using intermediate fusion achieved the best overall performance (AUC 0.93, accuracy 87.0%). Addition of mpMRI reduced performance (AUC 0.85, accuracy 83.0%). Decision curve analysis demonstrated positive net benefit across clinically relevant thresholds. Conclusion: XGBoost-based multimodal fusion improved preoperative BCR prediction following RP. DGC was the strongest individual predictor, while integration with PSMA-PET provided the best overall performance, supporting the potential of radiogenomic ML models for personalized risk stratification.

4
Point-of-Care Breath Volatile Organic Compound Analysis as a Tool for Lung Cancer Screening: A Pilot Feasibility Study

Pichkar, Y.; Manolakos, S.; Phillips, K. M.; Schabath, M. B.; Chaudhary, A.

2026-08-31 oncology 10.64898/2026.08.26.26361331 medRxiv
Top 0.1%
8.2%
Show abstract

Background: Low-dose computed tomography (LDCT) screening reduces lung cancer mortality but is limited by low uptake and associated with high rates of false-positives and indeterminate-nodules. Breath volatile organic compound (VOC) analysis is a non-invasive candidate biomarker approach that could complement LDCT, but prior work has relied on laboratory-based high-resolution mass spectrometry (HRMS), limiting point-of-care deployment. Methods: In this pilot study, breath samples were collected from 40 patients with treatment-naive, pathologically confirmed non-small cell lung cancer (NSCLC) and 25 lung-cancer-screening-eligible healthy controls. Paired samples were analyzed via a compact point-of-care GC-MS platform (CLARION) and a laboratory HRMS reference. Diagnostic classification models were built independently for each platform using elastic net logistic regression with leave-one-out cross-validation, and performance was evaluated by area under the receiver operating characteristic curve (AUC). Results: CLARION identified 103 VOCs across breath specimens, compared to over 900 identified by HRMS. Despite this difference in panel size, CLARION achieved diagnostic performance nearly identical to HRMS for distinguishing NSCLC cases from controls (AUC 0.864 vs. 0.863). Compared to controls, performance statistics were similar for early-stage NSCLC (AUC 0.854 vs. 0.841) and adenocarcinoma (AUC 0.770 vs. 0.787). VOCs of interest include p-cymene, phenol, propylbenzene, tetradecane, {beta}-ocimene, 2,3-dihydro-indole, and 1-methylthio-(Z)-1-propene. Conclusion: A compact, point-of-care breath GC-MS platform achieved diagnostic performance for NSCLC detection comparable to a laboratory HRMS reference despite a substantially smaller detected VOC panel. These findings support continued development of point-of-care breath VOC testing as a non-invasive, field-deployable complement to LDCT-based lung cancer screening.

5
Acquisition of Group B streptococcus colonization in preterm pregnancy

Bowers, A.; Elliott, J.; Book, N.; Krishna, S.; Hamburg-Shields, E.

2026-08-26 obstetrics and gynecology 10.64898/2026.08.21.26361025 medRxiv
Top 0.2%
7.4%
Show abstract

Objective: The purpose of this study was to estimate the negative predictive value (NPV) of screening for group B streptococcus (GBS) colonization in pregnant patients undergoing antepartum hospitalization for GBS colonization status at the time of preterm delivery. Study Design: This prospective, observational cohort study compared GBS colonization status upon initial hospital admission to that at the time of delivery. Pregnant patients at 22 to 35 weeks gestation admitted to the antepartum unit at a tertiary care hospital underwent standard screening for GBS colonization. When preterm labor progressed or iatrogenic preterm delivery was indicated, the GBS colonization test was repeated. Comparison of the sequential test results was performed to determine the NPV of the antepartum screening test for the intrapartum status. Results: 159 eligible patients were enrolled in the study, and 100 completed the study and were included in the analysis. The average gestational age at admission was 30 weeks 1 day (95% confidence interval [CI] 29w4d to 30w6d) and the average duration of pregnancy latency in the study group was 17.5 days (95% CI 15.1 to 19.8). GBS colonization rate at the time of admission was 18% and at the time of delivery was 20%. The NPV of GBS screening at admission was 91.5% (95% CI 83.2 to 96.5%) and the positive predictive value (PPV) was 72.2% (95% CI 46.4 to 90.3%). Conclusion: In a cohort of pregnant patients with preterm pregnancy complications, GBS screening at the time of antepartum hospital admission has an NPV of 91.5% (95% CI 83.2 to 96.5%) for GBS colonization at the time of preterm delivery. This is comparable to the published NPV of routine GBS screening for colonization status at term delivery.

6
Diagnostic accuracy of image-guided fine needle aspiration cytology in diagnosis of lung tumors

Bhandari, B.; Tiwari, M.; Adhikari, S.; Khanal, A.; Chettri, N. B.; Pandey, S.

2026-08-18 pathology 10.64898/2026.08.17.26360531 medRxiv
Top 0.2%
7.0%
Show abstract

Background: Lung cancer is leading cause of cancer related death globally. It is second most prevalent cancer among women worldwide and ranks third among females in Nepal. Contributing factors are smoking, tobacco use, air pollution, and delayed diagnosis. Image-guided fine needle aspiration cytology (FNAC) is rapid diagnostic technique for evaluating lung lesions. It is minimally invasive procedure with less complications. This study examine histocytologic makeup of lung lesions and link the results. Materials and Methods: This cross-sectional observational study included 65 patients irrespective of age and sex presenting with lung masses at Chitwan Medical College and Teaching Hospital from April 2023 to September 2024. After clinical and radiologic evaluation, all cases underwent image-guided FNAC and biopsy. Only specimens with unequivocal malignant features were classified positive. Histopathology served as diagnostic reference standard. Results: FNAC diagnosed 90.8% as malignant and 9.2% as benign. Biopsy confirmed malignancy in 92.3% of cases. FNAC demonstrated a sensitivity of 98.33%, specificity of 100%, positive predictive value(PPV) of 100%, and negative predictive value (NPV) of 83.33%. Concordance between FNAC and histopathological subtyping was 98.46%. Adenocarcinoma was most common subtype, followed by Squamous cell carcinoma(SCC) and small cell carcinoma. Smoking was most common contributing factor associated with malignancy. Conclusion and implications: Image-guided FNAC is an excellent diagnostic accuracy tool which possess higher level of concordance with biopsy in evaluating lung masses. It should be considered as frontline diagnostic tool, especially in resource limited settings. Keywords: FNAC, Lung cancer, Biopsy, SCC, Adenocarcinoma, Small cell carcinoma, Nepal

7
Report-Guided Semi-Supervised Learning for Scalable Prostate Cancer Detection on Biparametric MRI: Multicenter Prospective Validation and Multimodal Integration

Calado, A.; de Almeida, J. G.; Verde, A. S. C.; Tsiknakis, M.; Marias, K.; Regge, D.; Papanikolaou, N.; ProCAncer-I Consortium,

2026-08-07 radiology and imaging 10.64898/2026.08.05.26359781 medRxiv
Top 0.2%
6.8%
Show abstract

Purpose: To prospectively validate a semi-supervised learning framework with a lesion-only teacher model (RG-SSL-LOC) for scalable clinically significant prostate cancer detection on biparametric MRI (bpMRI) and assess its added value in multimodal models. Materials and Methods: A multicenter dataset of 13,706 bpMRI examinations (13,630 patients, 27 centers) was used for model development/validation. Three segmentation models (fully supervised learning [FSL], a state-of-the-art report-guided semi-supervised approach [RG-SSL], and the proposed RG-SSL-LOC) were evaluated at lesion- and case-level on external retrospective, external prospective, and internal prospective cohorts. Predictions from the best-performing model were combined with clinico-radiologic variables in a multimodal approach. All case-level results were compared with PI-RADS. Results: At lesion level, RG-SSL-LOC achieved higher median Dice than FSL and RG-SSL (0.49 vs 0.41 and 0.40; both p<.001). At case level, RG-SSL-LOC achieved area-under-the-curve (AUC) values of 0.83, 0.82, and 0.87 in the external retrospective, external prospective, and internal prospective cohorts, respectively. Compared with FSL, AUCs were 0.84 (p=.237), 0.80 (p=.020), and 0.84 (p<.001); compared with RG-SSL, AUCs were 0.83 (p=.929), 0.82 (p=.652), and 0.86 (p=.007); compared with PI-RADS, AUCs were 0.78 (p=.055), 0.83 (p=.652) and 0.86 (p=.480). Combined with clinico-radiological variables, RG-SSL-LOC significantly improved AUC versus clinico-radiological variables alone in the external retrospective (0.85 vs 0.80, p=.002), external prospective (0.87 vs 0.84, p=.008), and internal prospective (0.91 vs 0.88, p<.001) cohorts; in the latter, it reduced unnecessary biopsies by 15.19%. Conclusion: RG-SSL-LOC achieves better segmentation quality than other methods, demonstrates robust prospective multicenter performance and improves multimodal detection.

8
Cost Minimisation and Threshold Analysis of Anatomical Endoscopic Enucleation of the Prostate

Ong, J.; Lau, R.; Chow, K. M.; Huned, D.; Teo, R.; Lee, H. J.; Lim, E. J.; Aslim, E.; Lim, Y. W.; Chen, K.; Tan, Y. Q.; Park, J. J.; Tung, J.

2026-08-17 urology 10.64898/2026.08.15.26360519 medRxiv
Top 0.2%
6.8%
Show abstract

Introduction Anatomical endoscopic enucleation of the prostate (AEEP) techniques, including bipolar enucleation (B-TUEP), holmium laser enucleation (HoLEP), thulium laser enucleation (ThuLEP), and thulium fibre laser enucleation (ThuFLEP), demonstrate comparable clinical outcomes for benign prostatic hyperplasia. As clinical equivalence is increasingly established, cost becomes a key determinant of modality selection. We performed a cost minimisation analysis comparing index procedural costs across AEEP modalities from an institutional perspective. Methods A cost minimisation model was developed from the institutional perspective, incorporating amortised capital costs, maintenance, and consumables. In addition to the base-case scenario of 180 cases per year, we modelled two additional case volume scenarios: low (50 cases/year) and high (500 cases/year) volume. Thu:YAG laser fibres were modelled on two scenarios: disposable single-use, and reusable fibres (up to 10 cases per fibre). Breakeven analysis determined the threshold volume at which each laser modality achieves cost parity with B-TUEP, and one-way sensitivity analysis was performed on key cost parameters. Analysis was limited to index procedural costs calculated in Singapore dollars. Results At the base case of 180 cases per year, B-TUEP had the lowest index procedure cost (SGD 1,018), followed by ThuFLEP (SGD 1,584), ThuLEP (1,599), and HoLEP (SGD 1,655). Breakeven analysis demonstrated that HoLEP, ThuLEP, and ThuFLEP can never achieve cost parity with B-TUEP when laser fibres are single-use, as laser modalities carry higher costs on both capital and per-case dimensions. ThuLEP with reusable fibres (10 uses per fibre) was the only modality to cross below B-TUEP, at a breakeven volume of 198 cases per year. At 500 cases per year with reusable fibres, ThuLEP achieved the lowest cost (SGD 847), representing a 15.4% saving over B-TUEP. Sensitivity analysis identified annual case volume and B-TUEP loop cost as the most influential parameters. Conclusion Index procedural costs in AEEP are strongly influenced by case volume and consumable strategy. While B-TUEP remains cost-efficient at low volume, high-volume practice combined with reusable Thu:YAG fibre technology enables cost parity and potential cost advantage for laser enucleation. These findings highlight the importance of economies of scale and device utilisation in technology adoption.

9
Detection of intratumoral hypoxia in primary breast cancer using photoacoustic imaging

Shimizu, H.; Kawashima, M.; Kataoka, M.; Yoshikawa, A.; Asao, Y.; Takeuchi, Y.; Takada, M.; Saito, S.; Toi, M.; Masuda, N.

2026-08-23 oncology 10.64898/2026.08.20.26360033 medRxiv
Top 0.2%
6.5%
Show abstract

Background Tumor hypoxia and abnormal vasculature are closely associated with aggressiveness in solid tumors. Therefore, noninvasive assessment of these features in primary breast cancer is needed. Photoacoustic (PA) imaging is an emerging modality that enables real-time visualization of vascular architecture and hemoglobin oxygenation. Methods Breast PA imaging was performed in patients with primary breast cancer using a bed-type PA imaging system equipped with a hemispherical sensor and a flat specimen holder enabling mild breast compression. Three independent evaluators assessed predefined characteristics of tumor-associated vasculature: centripetal/disrupted vessels and intratumoral vessel-like signals. Oxygenation (S-factor) of tumor-associated vessels was estimated using dual-wavelength laser irradiation at 756 and 797 nm. Results PA imaging was performed in 9 tumors from 8 patients. Eight tumors were evaluable, after the exclusion of 1 tumor with segmental bloody discharge. Centripetal/disrupted vessels were identified in 7 tumors (87.5%). Intratumoral vessel-like signals were observed in all tumors (100%), with higher signal density than in surrounding tissue in 5 lesions (62.5%). Increased intratumoral signal density was associated with a higher Ki67-labeling index (two-sided P = .01). Mean intratumoral S-factor level (76.9% {+/-} 9.1%) was significantly lower than that of peritumoral vessels at 5 mm (86.4% {+/-} 5.9%) and 20 mm (88.5% {+/-} 4.9%) from the tumor margin (two-sided P < .01). Conclusion PA imaging with a flat specimen holder enables noninvasive visualization of tumor-associated vasculature with reduced oxygenation in primary breast cancer. This approach may provide a novel imaging platform for the early detection and functional assessment of breast cancer.

10
Maternal cell-free RNA versus combined screening for first-trimester prediction of early-onset preeclampsia: a nested case-control study

Satorres-Perez, E.; Castillo-Marco, N.; Igual, M.; Cordero, T.; Munoz-Blat, I.; Monfort-Ortiz, R.; Marcos-Puig, B.; Simon, C.; Garrido-Gomez, T.; Perales-Marin, A.

2026-09-02 obstetrics and gynecology 10.64898/2026.08.28.26361628 medRxiv
Top 0.2%
6.2%
Show abstract

Background. In Europe, first-trimester combined screening with the Fetal Medicine Foundation (FMF) algorithm identifies women at increased risk of preeclampsia who may benefit from personalized aspirin prophylaxis. However, a substantial proportion of early-onset preeclampsia (EOPE) remains undetected at clinically acceptable specificity. Objective. To evaluate the first-trimester performance of MaiRa for early-onset preeclampsia (EOPE) risk stratification by benchmarking it against FMF screening in the same women, characterizing discordant patient-level classification profiles and exploring potential implementation strategies. Study Design. This secondary case-control analysis was nested within the prospective, multicentre PREMOM cohort [NCT04990141], which enrolled women with singleton pregnancies across 14 tertiary hospitals in Spain. First-trimester MaiRa and FMF risk estimates were evaluated in the same 126 pregnant women, comprising 99 uncomplicated controls and 27 EOPE cases, defined by disease onset before 34 weeks. Discrimination was compared using a stratified paired bootstrap analysis of the areas under the receiver-operating-characteristic curves. Performance was assessed at prespecified clinical thresholds, and detection rates were evaluated at fixed false-positive rates. Universal and contingent MaiRa implementation strategies were also evaluated. Results. MaiRa showed greater first-trimester discrimination for EOPE than FMF combined screening (AUC, 0.974 vs 0.900; P=.040) and consistently achieved higher detection rates across fixed false-positive rates. At false-positive rates of 5% and 10%, MaiRa detected 85.2% and 92.6% of EOPE cases, compared with 44.4% and 70.4% for FMF, respectively. Patient-level analysis demonstrated that MaiRa identified 12 of 27 EOPE cases (44.4%) classified as low risk by FMF; these pregnancies generally exhibited less abnormal conventional first-trimester profiles, including fewer maternal risk factors, lower mean arterial pressure and lower uterine artery pulsatility index, yet 8 of 12 (66.7%) subsequently developed severe EOPE. Exploratory implementation analyses showed that universal MaiRa screening achieved the highest EOPE detection, whereas a contingent strategy using FMF for triage and reflex MaiRa testing reduced molecular testing to 35.7% of pregnancies while maintaining 77.8% sensitivity and 97.0% specificity. Conclusion. MaiRa provided greater first-trimester discrimination for EOPE than conventional combined screening and detected additional pregnancies that later developed severe disease despite less abnormal conventional screening profiles. The findings suggest that maternal plasma cfRNA profiling captures biological alterations not fully reflected by combined first-trimester screening and support further prospective evaluation in an independent, unselected obstetric population. Key words: early-onset preeclampsia; first-trimester screening; cell-free RNA; liquid biopsy; Fetal Medicine Foundation algorithm; combined screening; risk stratification; aspirin prophylaxis.

11
Assessment of impending pancreatic cancer in a cohort of new onset diabetes on basis of biomarker trajectory

Irajizad, E.; Lopez, C.; Chari, S.; Vykoukal, J.; Spencer, R.; Li, Y.; Dennison, J.; Koay, E.; McAllister, F.; Kim, M.; Young, M.; Hart, P.; Fischer, W.; Vandeneeden, S.; Wu, B.; Feng, Z.; Hanash, S.; Maitra, A.; Fahrmann, J.; Consortium for the Study of Chronic Pancreatitis, Diabetes, and Pancreatic Cancer (CPDPC),

2026-08-10 gastroenterology 10.64898/2026.08.06.26359908 medRxiv
Top 0.2%
5.6%
Show abstract

PURPOSE: To assess the predictive performance of panel protein biomarkers as well as an established algorithm that considers repeat biomarker testing for risk prediction of PDAC among a prospective cohort of patients with New-onset diabetes. PATIENTS AND METHODS: A panel of protein biomarkers (CA19-9, CA125, CEA, LRG1, REG3A and TIMP1) were assayed in 6,516 serially collected pre-diagnostic plasma samples from 2,121 NOD patients from the Consortium of Chronic Pancreatitis Diabetes and Pancreatic Cancer (CPDPC)-initiated NOD study who completed the 3-year study follow-up period. The specimen set included 25 pre-diagnostic samples from the 12 PDAC cases diagnosed during study follow-up. We applied a single threshold (ST) method, which considers biomarker levels at a single time point, as well as a previously established parametrical empirical Bayes (PEB) algorithm, which considers prior biomarker measurements, with case calls made based on pre-specified cutoffs corresponding to 1% 1-year risk. Resultant biomarker data as well as case calls were provided to the EDRN Data Management and Coordinating Center as part of a Prospective-sample-collection-Retrospective-Blinded-Evaluation (ProBE)-compliant Phase 3 biomarker validation study. Area under the Receiver Operating Characteristic Curves (AUC), sensitivity, specificity, population-level positive predictive value (PPV), and negative predictive value (NPV) are reported. RESULTS: The 3-year incidence of PDAC in the NOD cohort was 0.57%. When considering PDAC vs non-cancer controls, respective AUCs of individual protein biomarkers ranged from 0.52-0.94, with CA19-9 achieving the highest overall performance of 0.94 (95% CI: 0.86-1.00). At the pre-defined 1% 1-year risk threshold, CA19-9 yielded sensitivity of 83.3% at 97.2% specificity. Additional markers CEA, CA125, and TIMP1 demonstrated sensitivity of 33.3%, 41.7%, and 8.3%, respectively. In a subset of patients, CA19-9 first tested positive at a median (interquartile range [IQR]) of 7 months (4 to 14 months) prior to clinical PDAC diagnosis. Of the two PDAC cases missed by CA19-9 using the ST method, one (diagnosed with stage III PDAC) was detected using the PEBCA19-9 algorithm. CONCLUSION: In the setting of adult new onset diabetes, CA19-9 is a readily available and promising biomarker that can be leveraged for earlier detection of an underlying pancreatic cancer. Additional protein biomarkers may improve sensitivity for earlier detection of PDAC among cases with low CA19-9.

12
Prediction of Subsolid Pulmonary Nodule Evolution from Baseline CT Using Temporal Imaging Models

Bondarenko, M.; Qi, K.; Nowroozi, A.; Kim, J.; Kunzang, B.; Lee, A.; Liu, J.; Tran, N.; Weng, S.; Vella, M.; Chaudhari, G.; Schnizler, T.; Innanje, A.; Chen, T.; Sohn, J. H.

2026-08-13 radiology and imaging 10.64898/2026.08.12.26360292 medRxiv
Top 0.3%
5.6%
Show abstract

Background: Prediction of subsolid pulmonary nodule (SSN) progression from baseline CT may improve risk stratification and surveillance planning, but prior approaches have largely relied on fixed follow-up intervals. Methods: This retrospective single-center study evaluated interval-aware temporal imaging models for predicting future SSN growth and morphology across heterogeneous surveillance durations. A total of 24,946 longitudinal scan pairings derived from 2,543 clinician-reviewed SSNs in 426 patients were analyzed. A discriminative deep learning model predicted interval growth from baseline CT, segmentation masks, and interscan interval information, while a temporally conditioned generative model predicted future lesion morphology. Results: The discriminative model achieved an area under the receiver operating characteristic curve of 0.772 (95% confidence interval: 0.704-0.818), with sensitivity of 80.2% and specificity of 58.7% on the test cohort. The generative model predicted future lesion morphology with a Dice similarity coefficient of 0.706 +/-0.186. Prediction performance decreased with increasing follow-up duration, although both models generalized across intervals ranging from months to years. Conclusion: Interval-aware temporal imaging models enable the prediction of future SSN growth and morphology from baseline CT while accounting for variable surveillance intervals. These findings suggest a framework for time-aware, personalized risk assessment that may support individualized surveillance strategies and future AI-assisted management of pulmonary adenocarcinoma spectrum lesions.

13
Preoperative Prediction of Residual Cancer Burden After Neoadjuvant Chemotherapy in Breast Cancer: A Multimodal Machine Learning Approach and Implications for Clinical Decision Support

Dagdeviren, Y. K.; Semiz, H. S.; Inan, E. H.; Karakas, H. Y.; Durak, M. G.; Tezel, N.; Sevindik, M. C.; Kirmizibayrak, P. B.; Bekis, R.

2026-08-18 oncology 10.64898/2026.08.16.26360557 medRxiv
Top 0.3%
5.3%
Show abstract

Background. Residual cancer burden (RCB) after neoadjuvant chemotherapy (NAC) offers finer prognostic stratification than binary pathologic complete response, and increasingly guides adjuvant treatment intensity. Predicting four-tier RCB class from preoperative data could inform adjuvant planning before surgery, yet this remains an unmet need; and when two models reach equal discrimination, the key question is which generalizes most reliably. We compared a radiology-focused model with a fully integrated multimodal model for preoperative four-class RCB prediction. Methods. In a single-center, retrospective cohort of 328 patients treated with NAC followed by surgery, 64 clinicopathologic and radiologic variables were organized into thematic blocks. Two configurations were compared: a 17-variable radiology model (Model R) and a 62-variable multimodal model (Model ALL). Three algorithms (Random Forest, XGBoost, LightGBM) were evaluated with and without SMOTE using an 80/20 stratified split and 5-fold cross-validation. Model selection combined test AUC, macro-F1, cross-validation-to-test gap, nested cross-validation, bootstrap confidence intervals, and SHAP explainability, following the TRIPOD+AI guidance. Results. RCB classes were distributed as RCB-0 27.4% (n=90), RCB-I 10.4% (n=34), RCB-II 43.6% (n=143), and RCB-III 18.6% (n=61). Model R and Model ALL reached identical test AUC (0.838). Model ALL, however, achieved higher accuracy (0.636 vs 0.530) and macro-F1 (0.602 vs 0.598), together with a substantially smaller cross-validation-to-test gap (0.015 vs 0.099), pointing to more stable generalization; this gap difference persisted across all three algorithms. SHAP analysis showed that the multimodal model drew jointly on imaging phenotype, tumor biology, and disease extent. Both models remained weakest in the RCB-III class. Conclusions. At equivalent discrimination, the multimodal model was methodologically preferable for preoperative RCB prediction, owing to its stability and interpretability - qualities relevant to trustworthy clinical decision support. It remains investigational; a model flagging likely RCB-0 or RCB-III before surgery could prioritize adjuvant-therapy discussions earlier in the care pathway, pending prospective external validation.

14
CT ECV Mapper: an interactive 3D Slicer application with a batch-capable pipeline for voxelwise CT-derived extracellular volume mapping of the liver and hepatic tumors

Suzuki, M.

2026-08-11 radiology and imaging 10.64898/2026.08.09.26360018 medRxiv
Top 0.3%
5.2%
Show abstract

Background. Extracellular volume fraction (ECV) derived from contrast-enhanced CT is a validated marker of hepatic fibrosis and has been reported to differ between hepatocellular carcinoma (HCC) and intrahepatic cholangiocarcinoma. In published work it is obtained from a small number of hand-placed two-dimensional regions of interest, and the software that computes it is either tied to one manufacturer's workstation or based on spectral or dual-energy acquisition. We are not aware of an accessible tool that produces voxelwise liver ECV maps from conventional single-energy multiphase CT. Methods. We developed CT ECV Mapper, a scripted 3D Slicer extension with a three-layer architecture whose numerical core imports neither slicer nor vtk and is unit-tested outside 3D Slicer. The interactive application provides two-stage registration that the operator inspects and accepts before any ECV is computed, operator-placed three-dimensional regions of interest, user-adjustable calculation parameters, a voxelwise ECV color map and ROI statistics; the same logic layer can be driven unattended across a cohort. The tool was applied to the 164 patients of the public WAW-TACE multiphase HCC/TACE dataset that have both unenhanced and delayed-phase series. Results. 156 of 164 cases (95.1%) completed unattended. Whole-liver ECV had a median of 36.2% (interquartile range 31.9-41.5), consistent with published CT-ECV values for fibrotic and cirrhotic liver. Registering the arterial and portal phases on demand extended tumor ECV from the 38 lesions a conventional two-phase pipeline can reach to 248 lesions in 156 patients. Every failure was attributable to an identifiable mechanism: craniocaudal field-of-view mismatch between phases in six cases, aortic calcification within the blood-pool region in one, and in one case a labeling error in the source dataset, in which the series declared as unenhanced proved to be a second reconstruction of the portal venous phase; this was detected by the blood-pool validity check rather than by visual review. Conclusions. Voxelwise CT ECV mapping of the liver and of hepatic tumors is feasible from conventional multiphase CT on an open platform, both interactively and as an unattended batch, with quality-control instrumentation that fails explicitly and diagnosably. This is a technical development and feasibility report; the application has not been evaluated against a reference standard and no claim of clinical validity is made.

15
Clinical selectivity and failure modes of automated chest radiograph report evaluation metrics: a cross-dataset analysis of ReXErr-v1 and RadEvalX

Naidu, J.; Muralidharan, S.; Prashani, A.; Baskaradoss, V.

2026-08-12 radiology and imaging 10.64898/2026.08.10.26360043 medRxiv
Top 0.4%
4.2%
Show abstract

Objectives: To test whether radiology report evaluation metrics distinguish clinically meaningful errors from textual changes and align with radiologist-assessed error burden. Methods: Cross-dataset evaluation used ReXErr-v1 (2,708 report pairs; 5,724 paired error sentences) and 100 RadEvalX report pairs with expert error counts. BLEU-4, ROUGE-L and METEOR were assessed in ReXErr-v1; RadEvalX analyses included these plus BERTScore, CheXbert, RadGraph F1 and RadCliQ. Outcomes were ReXErr-v1 pairwise win rate and AUROC for clinical-content versus linguistic errors, and RadEvalX Spearman correlation with clinically significant error count and AUROC for any significant error. Confidence intervals used 10,000 clustered percentile bootstrap resamples; Holm adjustment-controlled multiplicity. Results: ReXErr-v1 paired-sentence win rates were 0.986 for BLEU-4, 0.999 for ROUGE-L and 0.998 for METEOR, but discrimination of clinical-content from linguistic errors was modest (AUROC 0.609-0.620). Penalty magnitude was strongly associated with textual change after adjustment for error type (normalised character edit distance coefficient 0.746; 95% CI 0.705-0.788; P<0.001). In RadEvalX, CheXbert showed the highest correlation with clinically significant errors (rho=0.413; 95% CI 0.223-0.578) and highest AUROC (0.742; 95% CI 0.638-0.836). Conclusions: Near-ceiling sensitivity to textual corruption did not imply sensitivity to clinical significance. CheXbert showed the highest alignment with expert error assessment, although pairwise superiority was not demonstrated over all comparators and performance remained moderate.

16
Internal and External Validation of an Ensemble Learning Model Integrating Zygote Morphokinetics with Conventional Embryo Assessment for Blastocyst Prediction

ZHAO, M.; LIU, J.; HAN, D.; ZHANG, C.; ZHOU, Y.; CHEN, S.; LIU, C.

2026-08-23 obstetrics and gynecology 10.64898/2026.08.19.26359526 medRxiv
Top 0.4%
4.1%
Show abstract

Objective: To perform internal and external validation of a gradient-boosted decision tree (GBDT) fusion model that integrates zygote morphokinetic parameters with conventional embryo assessment features for blastocyst prediction, and to compare its discriminative performance against senior embryologists. Methods: This retrospective cohort study included 631 normally fertilized zygotes from 218 treatment cycles. A GBDT fusion model integrating 84 zygote morphokinetic parameters and 8 conventional assessment features was evaluated internally (5-fold cross-validation) and externally on a public dataset of 523 embryos with blastocyst outcomes. Model performance was assessed using area under the ROC curve (AUC), area under the precision-recall curve (AUPRC), F1 score, sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV). Discrimination was compared with embryologist consensus using the DeLong test; agreement was assessed with Cohen's kappa. Results: The model achieved an internal AUC of 0.78 (95% CI 0.74-0.82), AUPRC 0.72, F1 0.73, sensitivity 0.74, specificity 0.77, PPV 0.72, and NPV 0.79. External validation on the public dataset demonstrated acceptable generalizability (AUC 0.76, 95% CI 0.71-0.81). The model significantly outperformed embryologist consensus (AUC 0.70, P<0.001) with moderate agreement (kappa=0.56). Decision curve analysis confirmed clinical net benefit at threshold probabilities of 0.15-0.55. Conclusions: The GBDT fusion model integrating zygote morphokinetics with conventional assessment demonstrates good discrimination and external generalizability for blastocyst prediction, providing an interpretable decision-support tool for embryo selection in IVF practice.

17
Mifepristone Priming with Misoprostol versus Intracervical Foley's Catheter with Misoprostol for Induction of Labour in Late Second and Third Trimester Intrauterine Fetal Death: A Prospective Comparative Study

Das, B.; Garg, P.

2026-08-10 obstetrics and gynecology 10.64898/2026.08.06.26359872 medRxiv
Top 0.4%
4.1%
Show abstract

Abstract Introduction Intrauterine fetal death (IUFD) beyond 24 weeks of gestation, particularly when accompanied by an unfavourable cervix, poses a distinct obstetric challenge in achieving safe and timely vaginal delivery while minimising maternal distress. Mifepristone priming followed by misoprostol and intracervical Foley's catheter combined with misoprostol are both established approaches for cervical ripening and induction of labour in this setting, but direct comparative data especially from Indian tertiary care populations remain limited. Methods This prospective comparative study was conducted in the Department of Obstetrics and Gynaecology, Kamla Raja Hospital, Gajra Raja Medical College (GRMC), Gwalior, Madhya Pradesh, India, over a two-year period (November 2020 - October 2022). One hundred and fourteen women with ultrasonography-confirmed IUFD beyond 24 weeks of gestation were alternately allocated to Group A (n=57; oral mifepristone 200 mg followed by gestational-age-adjusted vaginal misoprostol) or Group B (n=57; intracervical 16F Foley's catheter followed by gestational-age-adjusted vaginal misoprostol). Outcomes assessed included pre- and post-induction Bishop score, induction-to-delivery interval, misoprostol dose requirement, need for oxytocin augmentation, mode of delivery, blood loss, maternal complications, pain (visual analogue scale, VAS), and patient satisfaction. Results Baseline age, parity, gestational age, and pre-induction Bishop score were comparable between groups (p>0.05). The mean post-induction (24-hour) Bishop score was significantly higher in Group A (7.39+/- 2.07) than Group B (6.37+/-1.89; p=0.007). The mean induction-to-delivery interval was significantly shorter in Group A (25.43+/- 6.84 hours) than Group B (29.26+/- 5.54 hours; p=0.0014), and the median misoprostol dose requirement was significantly lower in Group A (50 mcg) than Group B (100 mcg; p<0.01). Mode of delivery, blood loss, oxytocin augmentation requirement, and overall maternal complication rates did not differ significantly between groups (all p>0.05). Pain scores were significantly lower in Group A (VAS 2.83+/- 1.16) than Group B (VAS 6.18+/- 1.69; p<0.0001), while patient satisfaction was comparable between groups (96.5% vs. 91.23%; p=0.244). Conclusions Both mifepristone-misoprostol and Foley's catheter-misoprostol regimens are safe and effective methods for induction of labour following IUFD beyond 24 weeks of gestation with an unfavourable cervix. Mifepristone priming achieved a shorter induction-to-delivery interval, lower total misoprostol requirement, and substantially less procedural pain, making it an attractive first-line option where available, while Foley's catheter remains a safe, low-cost, and widely accessible alternative, notwithstanding lower patient comfort.

18
Initial Staging 18F-FDG PET/CT for Coronary Artery Calcium Scoring to Assess Cardiovascular Risk in Women with Breast Cancer

Fleming, M. R.; Tayon, K. G.; Schneider, A.; McPherson, A. D.; Bianco, S. M.; Parent, E. E.; Sharma, A.; Lin, G.; Norton, N.; Ray, J. C.

2026-08-26 cardiovascular medicine 10.64898/2026.08.24.26361275 medRxiv
Top 0.4%
4.0%
Show abstract

Background. Cardiovascular disease is a leading cause of death among women with breast cancer, and the 2026 ACC/AHA dyslipidemia guideline endorses coronary artery calcium (CAC) scoring to guide statin therapy before cardiotoxic treatment. Breast cancer patients routinely undergo staging 18F-fluorodeoxyglucose PET/CT, whose low-dose CT visualizes the coronary arteries, thus enabling CAC quantification at no additional cost or radiation. Methods. In this single-center retrospective study, consecutive women with newly diagnosed breast cancer undergoing staging 18F-FDG PET/CT (2009?2021) had semi-automated Agatston CAC scoring performed on the low-dose CT and were stratified by CAC presence (CAC-P) versus absence (CAC-A). We assessed a composite of cardiac diagnostic testing (stress testing, coronary CT angiography, invasive angiography), clinical events, and reclassification of statin eligibility per ACC/AHA guideline thresholds in a prevention-eligible subgroup. Results. Among 276 women (mean age 55.5 years; median follow-up 7.1 years), CAC was present in 68 (25%) but was clinically reported in only 5.4%. CAC-P was associated with more cardiac testing (34% vs 12%; age-adjusted hazard ratio 2.75, 95% CI 1.43?5.28) and, though underpowered, with more atherosclerotic events (7.4% vs 1.4%), but not with the all-cause composite. In the prevention-eligible subgroup (n=39), CAC scoring would have changed statin eligibility in 64%, initiating therapy in 62% of CAC-P women and supporting de-prescribing in 67% of CAC-A women. Conclusions. CAC can be feasibly quantified from staging PET/CT in women with breast cancer and would frequently reclassify statin eligibility at no additional cost or radiation, yet is rarely reported.

19
Evaluation of risk stratification at presentation using the Alinity high-sensitivity cardiac troponin I assay

Li, Z.; Fujisawa, T.; Skadberg, O.; Fineran, P.; Thurston, A. J.; Tew, Y. Y.; Aakre, K. M.; Mills, N. L.; Wereski, R.; the POC-ET Investigators,

2026-08-31 cardiovascular medicine 10.64898/2026.08.29.26361405 medRxiv
Top 0.4%
4.0%
Show abstract

Background: High-sensitivity cardiac troponin (hs-cTn) assays enable safe early discharge of patients at very low risk for myocardial infarction. We previously developed a single-sample rule-out pathway using the ARCHITECT hs-cTnI assay to risk stratify patients with suspected acute coronary syndrome. In a secondary analysis of the POC-ET (Point of Care Evaluation of High-sensitivity Cardiac Troponin) study, we evaluated performance of risk stratification with the Alinity hs-cTnI assay. Methods: Patients presenting with possible myocardial infarction in the POC-ET (NCT05665127) study were included. The primary outcome was type 1, 4b or 4c myocardial infarction or cardiac death at 30 days. Cardiac troponin I (cTnI) was measured in stored materials using the ARCHITECT and Alinity hs-cTnI assays. The sex-specific 99th percentile upper reference limit (URL) are 34 ng/L in men and 16 ng/L in women for both assays. Agreement was assessed with Bland-and-Altman limit of agreement method, Passing Bablok regression, and Pearson's correlation coefficient. Distributions of presentation measurements were compared with Kolmogorov-Smirnov test. Performance was evaluated in the overall population and prespecified subgroups. The negative predictive value (NPV) and sensitivity were determined and proportion of patients identified as low, intermediate, and high risk were calculated and modelled using ordinal logistic regression. Results: In 986 patients (60 [51-70] years, 38% female), 78 (7.9%) had a primary outcome. Strong agreement was found in the raw cTnI measurements (99% samples within the Bland-Altman limit of agreement; correlation coefficient r: 0.967 (95% CI 0.964-0.969, P<0.001); Passing Bablok regression: slope 1.12 [1.11-1.13], intercept -0.16 [-0.18 to -0.13]). At presentation, distributions of cTnI measurements by the two assays were similar (P=0.810). Both assays showed comparable diagnostic performance using a risk stratification threshold of <5 ng/L and the sex-specific diagnostic threshold, with the same NPV (Alinity 100 [99.7-100]% versus ARCHITECT 100 [99.7-100]%) and sensitivity (Alinity 100 [97.3-100]% versus ARCHITECT 100 [97.3-100]%). Similar proportions of patients stratified as low- (Alinity 67% versus ARCHITECT 67%), intermediate-risk (23% versus 24%) and high-risk (10% versus 9%) at presentation with minor reclassification. Similar efficacy was observed across subgroups stratified by sex, age, history of myocardial infarction, renal function, and symptom duration. Conclusions: The Alinity hs-cTnI and the ARCHITECT hs-cTnI assays can be used interchangeably in the assessment of suspected myocardial infarction with comparable safety and efficacy.

20
Progesterone and hCG in expectant management success in tubal ectopic pregnancy: retrospective single-centre cohort study

Ahmad, A. K.; Pandrich, M.; Naik, A.; Astruc, A.; Lafferty, K.; Shah, N. M.; Ofili-Yebovi, D.

2026-08-07 obstetrics and gynecology 10.64898/2026.08.05.26359789 medRxiv
Top 0.5%
3.5%
Show abstract

Background: Early access to pregnancy assessment units now detects many tubal ectopic pregnancies (TEP) at a stage when they could resolve spontaneously, creating a management dilemma. Methods: We performed a hypothesis-generating exploratory analysis in a retrospective study to assess whether serum progesterone (P4) levels in women with TEP are associated with management outcome. Results: Ninety-one cases of TEP managed in a single centre over three years were analysed. Receiver operating characteristic (ROC) curve analysis was used to explore serum levels of progesterone (P4), first human chorionic gonadotropin (hCG) and peak hCG (alone and in combination) in relation with successful completion of expectant management. Decision-tree analysis using first hCG and P4 was additionally performed to explore clinical sequential risk stratification. 23% (n=21) successfully completed expectant management. P4 concentrations in the expectant management group (median 3 nmol/L, IQR 2.00 to 8.50) were significantly lower than in those requiring surgical or medical management (median 17 nmol/L, IQR 5.75 to 29.25; p=0.0002). Area under the ROC curve (AUC) values for P4, log10 first hCG, log10 peak hCG and P4 with log10 first hCG were 0.766, 0.814, 0.811 and 0.835, respectively, for predicting successful expectant management. However, hCG was not significantly outperformed. Nonetheless, Youden optimised thresholds for hCG and P4 are reported, alongside decision-tree analysis that identified sequential first hCG and P4 thresholds associated with successful expectant management. Conclusion: Lower P4 levels are associated with successful expectant management of TEP but they do not outperform hCG either alone or as an adjunctive marker.