Diagnostics
○ MDPI AG
Preprints posted in the last 90 days, ranked by how well they match Diagnostics's content profile, based on 50 papers previously published here. The average preprint has a 0.08% match score for this journal, so anything above that is already an above-average fit.
Ahmed, A. F. F.
Show abstract
Background Lung cancer mortality is rising in Libya, but access to molecular diagnostics for EGFR mutations--essential for guiding tyrosine kinase inhibitor therapy--remains severely limited. Selecting an appropriate testing platform requires balancing analytical performance against cost and infrastructure constraints. Methods We conducted a prospective comparative validation study using formalin-fixed paraffin-embedded (FFPE) tissue samples from Libyan non-small cell lung cancer (NSCLC) patients. Following stringent DNA quality control, samples were tested in parallel across four platforms: multiplex real-time PCR (MRT-PCR), reverse hybridization strip assay (RHSA), agarose gel electrophoresis (AGE), and immunohistochemistry (IHC). Performance was assessed by inter-method concordance, turnaround time, and cost per test. Results Of 30 initial samples, only six (20%) met quality thresholds (A260/A280 1.70-1.90; concentration [≥]10 ng/{micro}L), highlighting pre-analytical challenges. Three samples harbored EGFR exon 19 deletions. A critical discordance was identified: one sample tested negative by MRT-PCR (Ct {approx}38, {Delta}Ct=13) but positive by RHSA, AGE, and IHC, indicating a false-negative result from the reference method. IHC and RHSA offered the most favorable balance of cost (USD 40-75/test) and operational feasibility, while MRT-PCR (USD 150/test) required specialized infrastructure. Conclusions Relying solely on automated PCR may lead to under-diagnosis in low-cellularity or degraded FFPE samples. We recommend a hybrid algorithm: IHC as a cost-effective primary screen, followed by RHSA for confirmation. This approach optimizes resource allocation and improves diagnostic equity in Libya.
stern, N.
Show abstract
**Background:** Acute pancreatitis (AP) is a common gastrointestinal emergency with a subset of patients progressing to severe acute pancreatitis (SAP), which carries substantial morbidity and mortality. Current clinical severity scores such as BISAP, APACHE II, Ranson, and the Modified CT Severity Index require upon 48 hours of observation before reliable assessment is possible, limiting early triage. Machine learning (ML) approaches using routine admission laboratory values may enable earlier, more accurate prediction. **Methods:** We evaluated 11 models spanning three architectural families classical ML (Logistic Regression, Random Forest, Gradient Boosting), feedforward deep learning (MLP, Residual MLP, Attention MLP), and recurrent deep learning (LSTM, Stacked LSTM, Bidirectional LSTM, LSTM+Attention, CNN-LSTM) on a Chinese AP cohort of 722 patients (585 severe, 137 mild) labelled according to the 2012 Revised Atlanta Classification. Performance was assessed via 5-fold stratified cross-validation using AUC-ROC, F1 score, sensitivity, specificity, and PPV, with decision thresholds optimised for maximal F1. **Results:** Random Forest achieved the highest AUC of 0.877 (F1=0.917, sensitivity=96.8%, PPV=87.1%), followed closely by Gradient Boosting (AUC=0.874, F1=0.918). Classical ML models consistently outperformed deep learning counterparts. CNN-LSTM was the best recurrent model (AUC=0.777) but remained inferior to all classical approaches. LSTM-family models produced AUC values of 0.684-0.777, reflecting the cross-sectional tabular nature of the data. **Conclusions:** Random Forest provides robust, high-sensitivity early prediction of SAP severity using routine admission data. External prospective validation is required before clinical deployment. **Keywords:** acute pancreatitis; severity prediction; machine learning; random forest; deep learning; LSTM; Revised Atlanta Classification; early triage
Takeuchi, T.; Nomiya, A.
Show abstract
Background: A 2019 report from our institution described a multilayer artificial neural network (ANN) for predicting prostate cancer at biopsy in 334 patients, trained with TensorFlow 1.x and evaluated at three fixed step counts without separating hyperparameter selection from test evaluation. We re-analyzed an expanded cohort from the same institution using contemporary machine-learning practice. Methods: We pooled all available biopsy episodes from the same institutional database (n = 526; 524 after excluding one non-binary outcome code and one record with missing digital rectal examination [DRE] data), retaining the same seven predictors used in the original report (age, prior biopsy history, PSA, prostate volume, DRE, and MRI diffusion-weighted imaging findings in the peripheral and transition zones). Because 27 patients contributed more than one biopsy episode, we used patient-ID-grouped, stratified k-fold cross-validation (StratifiedGroupKFold; scikit-learn 1.8.0) with 3 and 5 folds, repeated over 10 random partitions, to avoid leakage between folds. Four classifiers were compared: L2-regularized logistic regression, gradient boosting, random forest, and a shallow (single hidden layer) multilayer perceptron. Two outcomes were modeled: detection of any prostate cancer, and detection of clinically significant prostate cancer (Gleason score [≥] 7). Results: Any-cancer prevalence was 55.7% (292/524) and Gleason score [≥] 7 prevalence was 39.7% (208/524). With repeated 5-fold cross-validation, gradient boosting gave the highest discrimination for any prostate cancer (mean AUC 0.826, 95% CI 0.823-0.830) and for Gleason score [≥] 7 (mean AUC 0.855, 95% CI 0.852-0.859), closely followed by random forest and logistic regression (AUC 0.81-0.85). The shallow multilayer perceptron performed worse and less consistently than the other three models (any-cancer AUC 0.671; Gleason score [≥] 7 AUC 0.742) and than the deeper five-hidden-layer ANN reported in 2019. Results with 3-fold cross-validation were essentially unchanged. Conclusions: In an expanded cohort, regularized logistic regression, gradient boosting, and random forest all discriminated prostate cancer at biopsy at least as well as the previously reported multilayer ANN, using far simpler models and a methodology that separates hyperparameter tuning from performance estimation. A shallow neural network offered no advantage over these simpler alternatives in this sample size. This is a preprint; the study has not undergone external peer review.
Scotland, K. B.; Ojo, O. A.; Anokwuru, F.; Javaherforoush, J.; Jimenez, J.; Teoh, A.; Suryavanshi, B.; Chan, R.
Show abstract
Background: Calcium oxalate nephrolithiasis is the most common type of kidney stone disease. Dietary oxalate intake is an important modifiable factor. Assessing dietary oxalate exposure in clinical practice poses challenges due to limitations of traditional dietary recall tools and variability in food composition data. Artificial intelligence (AI) applications in mobile health may offer scalable solutions for better dietary monitoring and kidney stone prevention. We examined the ability of StoneFree AI to estimate dietary oxalate from verbal and image-based food inputs. Objective: To evaluate the accuracy and limitations of StoneFree AI, for estimating dietary oxalate intake from verbal food descriptions and meal images, and to evaluate errors from entries that may inform future clinical use in kidney stone prevention. Methods: StoneFree AI is a cross-platform mobile application that uses a multimodal large language model (Google Gemini) to interpret verbal food descriptions and visual food images. The identified foods were mapped to oxalate values using the Harvard Oxalate Database. System performance was evaluated using 804 verbal food entries and 276 portion-size food images obtained from the ASA24 dietary assessment database. Verbal inputs were compared with reference oxalate values using absolute error and predefined agreement thresholds ({+/-}1, {+/-}5, {+/-}10 mg). Image-based inputs were evaluated against mutually exclusive primary error categories, including food identification, portion estimation, ingredient recognition, oxalate reference selection, and non-analyzable cases. Results: For verbal food entries, the AI system showed strong agreement with reference oxalate values. Overall, 82.1% of estimates were within {+/-}1 mg, 91.5% within {+/-}5 mg, and 94.5% within {+/-}10 mg of reference values. The mean absolute error was 3.32 mg, the median absolute error was 0.10 mg, and the concordance correlation coefficient (CCC) was 0.860. Image-based inputs showed a higher overall error rate of 63.0%, primarily due to food identification errors (33.0%), inaccurate portion estimation (11.0%), and ingredient recognition errors (9.8%). Most errors occurred with visually complex meals, such as mixed dishes and grain-based foods. Conclusions: AI-assisted estimation of dietary oxalate intake demonstrated high accuracy when structured verbal inputs were used but was less reliable for image-based meal analysis. These findings suggest AI-enabled mobile tools may support dietary monitoring for kidney stone prevention, particularly when user input is structured. Further refinement of computer vision models and prospective clinical validation are required before widespread clinical implementation.
Roy, S.; Soroar, M. K. I.; Ara, H.; Nur, S. A.; Akanda, R. A.; Saha, S.; Alam, M. M.
Show abstract
Background with objective: Detecting EGFR mutations is critical for treating lung adenocarcinoma with highly effective targeted therapies. However, standard genetic testing is expensive, complex, and often unavailable in resource-limited settings like Bangladesh. Because elevated serum CEA has been linked to these genetic alterations, it could serve as an accessible screening tool. This study aims to evaluate the association between serum CEA levels and EGFR mutation status to determine if routine CEA testing can reliably predict these mutations and guide treatment. Methodology: In this cross-sectional analytical study, we recruited 58 patients with histologically confirmed treatment naive lung adenocarcinoma. The presence of EGFR mutations in the ctDNA was determined via ARMS (Amplification Refractory Mutation System) PCR. Patient data was statistically analyzed to assess the diagnostic correlation between serum CEA levels and the presence of EGFR mutations. Result: The overall EGFR mutation rate was 43.1% with exon 19 deletion (48%) and exon 21 mutations (44%) were the predominant types. Median serum CEA levels were significantly higher in patients with EGFR mutations compared to wild-type cases (14.6 ng/ml vs 2.8 ng/ml, p<0.001). A multivariate analysis revealed a 14% increased likelihood of an EGFR mutation for 1 ng/ml rise in serum CEA. Furthermore, serum CEA showed strong diagnostic accuracy for ctDNA samples at a 6.39 ng/ml cut-off (AUC 0.82, sensitivity 68.0%, specificity 84.8%). Conclusion: Serum CEA is a valuable, cost-effective, and non-invasive biomarker demonstrating significantly higher levels and strong diagnostic accuracy in EGFR-mutated lung adenocarcinoma compared to wild-type cases.
Salis, F.; Tallone, N.; Fina, P. R.; Massobrio, R.; Bellacosa Marotti, R.; Conti, D.; Fuso, L.; Mariani, L.; Ferrero, A. M.; Accomasso, F.; Arena, A.; Borella, F.; Casula, V.; Cosma, S.; De Grandis, T.; Grisaru, D.; Lacalandra, A.; Pereira Sanchez, A.; Seracchioli, R.; Robba, E.; Roccio, M.; Gerace, F.
Show abstract
Ovarian cancer is recognized as the deadliest gynecological malignancy. Diagnosis at advanced stages and the lack of effective screening program lead to poor survival rates, dropping to 17-39 % in stage III-IV diseases. Ultrasound (US) is the primary imaging modality for ovarian structures evaluation, but it is strongly affected by the operator expertise due to the complexity of adnexal masses and the physiological variability of ovarian morphology throughout a womans lifecycle. The International Ovarian Tumor Analysis (IOTA) group introduced definitions and predictive tools to standardize gynecological US interpretation. However, these tools still rely on subjective interpretation, thus highlighting the need for more objective solutions. Recent studies have explored artificial intelligence (AI) algorithms for gynecological US, mainly focusing on adnexal masses classification. Conversely, a robust solution supporting the identification and description of healthy and tumoral ovarian structures is still lacking. This paper proposes OvAi Focus, a framework including (i) a segmentation module for the identification of healthy ovaries, plus solid and cystic components of adnexal masses; (ii) a morphology module for the extraction of IOTA-based keywords to describe adnexal masses morphology. Segmentation module results were compared to ground truth masks, showing DICE scores from 0.62 for functional ovary to 0.87 for the whole adnexal mass. Morphology module was tested through interobserver agreement analysis, obtaining Fleiss Kappa from 0.16 to 0.57 and Percent Agreement from 47 to 90 %, in line with existing literature. OvAi Focus represents an innovative solution which could help overcoming subjectivity in gynecological US imaging interpretation.
Chanian, R.; Mishra, D.; Jain, R.; Sharma, N.; Khurana, A.; Tripathi, R.; Tripathi, A.; group, G.-I. s.; Wadhwa, N.; Noble, J. A.; Thiruvengadam, R.; Desiraju, B. K.; Bhatnagar, S.
Show abstract
Preterm birth is the leading cause of neonatal death. Despite sustained efforts to identify high-risk women in the mid-trimester, accurate prediction remains difficult. Quantitative cervical ultrasound texture has been proposed as a predictor of spontaneous preterm birth. However, earlier models were developed in small single-centre samples and were not externally validated. We developed image-texture (Local Binary Patterns with a Random Forest), deep-learning (Vision Transformer), clinical-variable, and multimodal models to predict spontaneous preterm birth on the prospective GARBH-Ini cohort. We then externally validated our best models on an independent cohort scanned on a different ultrasound machine. Our best overall model reached an internal-test area under the receiver-operating-characteristic curve of 0.71 (95% CI 0.60, 0.82), but performed modestly at 0.52 (95% CI 0.38, 0.64) externally. The deep-learning and multimodal models did not perform better. Discrimination appeared higher in a clinically high-risk subgroup at the 34-week threshold. These estimates were imprecise because of few cases and need to be confirmed in future studies. Among the several likely reasons for the modest external performance is the heterogeneity of preterm birth. Predicting distinct preterm-birth subtypes separately, and integrating additional biomarkers and data domains, might improve model performance. Keywords: preterm birth; cervical ultrasound; prediction model; external validation; deep learning
Sholklapper, T. N.; Li, M.; Srivastava, A.; Wagh, A.; Handorf, E.; Beck, J. R.; Abbosh, P.
Show abstract
Importance There is growing interest to avoid radical cystectomy (RC) in patients with muscle-invasive bladder cancer (MIBC) who receive neoadjuvant chemotherapy and achieve pathological complete response (ypCR). To achieve this goal, molecular biomarkers will likely need to be used to enhance clinical staging given the limitations of evaluation by cystoscopy, cytology, and cross-sectional imaging. There are no studies evaluating whether safe RC avoidance (SaRCA) would be cost effective and what impacts it would have on quality of life (QoL) and survival. Objective This study models the potential economic, QoL, and survival costs/benefits of a ypCR biomarker as it relates to SaRCA using a decision analysis and Markov Model (MM). Methods/Materials A decision tree and MM was created to compare the expected costs of initial treatment, QoL, and survival under one strategy where all patients undergo RC after neoadjuvant treatment versus an alternative strategy where all patients would be subjected to the biomarker test with biomarker-positive patients (those with presumed residual disease) undergoing RC, while biomarker-negative patients (presumed complete responders) would undergo surveillance for up to 20 years. ypCR rates to neoadjuvant therapy, survival with and without RC, quality adjusted life years (QALY), and costs were abstracted from the literature. Test cost, sensitivity, and specificity were also abstracted from the literature for multiple clinical or liquid biopsy approaches. Results Broadly, SaRCA approaches are cost effective with the exception of systematic endoscopic evaluation (SEE). All testing approaches result in higher QALY and overall life expectancy compared to no testing. The cost of the test is offset by decreased usage of RC to realize a cost savings. These domains are further improved when cisplatin-based chemotherapy is replaced with emerging neoadjuvant therapies. Conclusions and relevance Modeling supports the development of accurate biomarker tests which can distinguish residual disease states to enable SaRCA. Such a biomarker could be used to avoid an expensive and risky operation, and unexpectedly would provide a survival benefit by reducing the number of perioperative mortalities in patients achieving ypCR. Development of an accurate biomarker-based test is likely to reduce cost and increase QoL and survival. An accurate biomarker test would have utility for patients, payers, hospitals, and physicians.
Kumar, D.; Kapoor, S.; Gowda, A.; Gupta, D.; Mittal, H.; Sood, S.
Show abstract
Background: Non-invasive hemoglobin measurement offers a painless and rapid alternative to conventional blood-based testing. The Vivaray Hb pro is a handheld photoplethysmography-based device designed for point-of-care hemoglobin assessment without blood sampling. We evaluated the clinical performance of the Vivaray Hb pro by comparing device-generated hemoglobin values with those obtained from a calibrated laboratory blood cell counter. Methods: In this cross-sectional, non-randomized clinical performance study, participants aged [≥]8 years were prospectively recruited. Hemoglobin was measured non-invasively using the Vivaray Hb pro and compared with venous blood samples analyzed on a calibrated Coulter counter. Agreement between methods was assessed using Bland-Altman analysis, including regression-based evaluation for proportional bias. Mean absolute error (MAE) and proportions of measurements within tolerance limits were also calculated. Complete paired measurements were available for 763 individuals. Results: Bland-Altman analysis demonstrated hemoglobin-dependent bias, with overestimation at lower hemoglobin levels and underestimation at higher levels. Regression-based analysis showed proportional bias ({beta}? = -0.178), indicating decreasing difference with increasing hemoglobin concentration. The MAE was 1.5 g/dL but was lower (1.2) in the clinically predominant ranges of 8.1-13 g/dl. Comment: The results support the use of the Vivaray Hb pro as a noninvasive hemoglobin screening and point-of-care assessment tool, particularly in settings where rapid, painless, and repeatable measurements are desirable.
Reddy Chimmula, R.; Yong, C.; Love, H. L.; Shiradkar, R.; Holmes, J.; Nair, V.; Tann, M.; Bahler, C.; Oderinde, O. M.
Show abstract
Background: Biochemical recurrence (BCR) occurs in up to 40% of men following radical prostatectomy (RP). Current risk models rely primarily on clinicopathologic variables and may not fully capture the biological heterogeneity associated with recurrence. The Decipher Genomic Classifier (DGC), prostate-specific membrane antigen positron emission tomography (PSMA-PET), and multiparametric magnetic resonance imaging (mpMRI) provide complementary prognostic information that may improve prediction. Objective: To develop and evaluate machine learning (ML) models integrating DGC, PSMA-PET, and mpMRI for preoperative prediction of BCR following RP. Methods: This retrospective study included patients with available preoperative DGC, PSMA-PET, mpMRI, and clinicopathologic data. Logistic regression (LR), random forest (RF), and XGBoost models were developed using single- and multimodality feature combinations. Early- and intermediate-fusion strategies were evaluated. Performance was assessed using an area under the receiver operating characteristic curve (AUC) and accuracy. Clinical utility was evaluated using decision curve analysis. Results: XGBoost consistently outperformed LR and RF. DGC achieved the highest single-modality performance (AUC 0.94, accuracy 86.7%). Among multimodal models, DGC combined with PSMA-PET using intermediate fusion achieved the best overall performance (AUC 0.93, accuracy 87.0%). Addition of mpMRI reduced performance (AUC 0.85, accuracy 83.0%). Decision curve analysis demonstrated positive net benefit across clinically relevant thresholds. Conclusion: XGBoost-based multimodal fusion improved preoperative BCR prediction following RP. DGC was the strongest individual predictor, while integration with PSMA-PET provided the best overall performance, supporting the potential of radiogenomic ML models for personalized risk stratification.
Dewi, Y. K.; Chudori, Y. N.
Show abstract
Reliable DNA isolation is a critical prerequisite for PCR-based food authentication, particularly for meat products where complex matrices may compromise DNA quality and amplification efficiency. This study aimed to analytically validate an automated DNA extraction method from meat matrices using Qiagen QIAcube Connect in combination with the DNeasy(R) Mericon Food Kit. Validation parameters included DNA concentration, total yield, purity, integrity, and assessment of PCR inhibitors using real-time PCR targeting the porcine cytochrome b gene. The method produced a mean DNA concentration of 219.5 ng/{micro}L with an average yield of 21,519.7 ng, exceeding predefined acceptance criteria. Agarose gel electrophoresis confirmed DNA fragment sizes larger than the target amplicon, indicating suitability for PCR analysis. Real-time PCR evaluation demonstrated excellent linearity (R2 = 0.99-1.00), amplification efficiencies between 90.34% and 99.84%, and mean {Delta}Ct values of 0.10, confirming the absence of PCR inhibition. These results indicate that the validated automated method is robust, reproducible, and suitable for routine PCR-based meat species authentication in food control laboratories.
Uskova, N. G.; Gombolevskiy, V. A.; Chernina, V. Y.; Burenchev, D. V.; Akhaladze, D. G.; Panina, E. V.; Karachunskiy, A. I.; Tereschenko, G. V.; Goncharov, M. Y.; Soboleva, E. A.; Konopleva, E. I.; Bydanov, O. I.; Plekhov, S. Y.; Grachev, N. S.
Show abstract
Background. Lung metastases in osteosarcoma (OS) are the main cause of the death. The accuracy of the diagnosis of nodules by computed tomography (CT) of the lungs is critically important for determining the disseminated stage of the disease and planning surgical treatment. The use of artificial intelligence (AI) in the search for lung nodules increases the accuracy of diagnosis and reduces the chance of missing metastases. Objective: to evaluate the accuracy of lung nodules diagnosis in adolescents with OS using AI. Methods. A retrospective assessment of CT scans of adolescents with OS was performed. A pathological nodule with an average size of [≥]4 mm was considered a target finding. The diagnostic accuracy of an AI algorithm previously trained on an adult dataset was evaluated, and the number of false positives (FP) and false negatives (FN) was determined. Sensitivity, specificity, accuracy, area under the ROC curve (AUC), positive predictive value, negative predictive value, and F1-measure were calculated. Based on the obtained results, the effectiveness of the algorithm was assessed. Results. 248 CT scans of adolescents with OS were evaluated. The following results were obtained: in 5 cases, the AI algorithm showed a FP result (2.02%), in 34 cases, it showed a FN result (13.71%), and in 209 cases, a correct result (both true positive and true negative) (84.27%). The diagnostic accuracy of the algorithm was 0.843 (95% CI 0.794-0.887). The application of the AI algorithm in the practice of an X-ray doctor in a specific clinical task would allow to increase the sensitivity from 0.805 to 0.891, while ensuring an absolute decrease in the number of FN results by 8.59% and a relative decrease by 44%. Conclusion. The obtained results confirm the practical value of the application of the AI algorithm and justify the implementation of AI-assisted systems in the diagnostic protocols for lung metastases in adolescents with OS.
Isakov, V.; Goncharov, A.; Israpilov, M.
Show abstract
Background and Aims: Vibration-controlled transient elastography (VCTE) and two-dimensional shear-wave elastography (2D-SWE) are used to assess liver stiffness in metabolic dysfunction-associated steatotic liver disease (MASLD); patients may be assessed by different methods over time. Because 2D-SWE fibrosis cut-offs are not uniformly validated and no biopsy or magnetic resonance elastography reference is available, we asked whether switching methods moves patients across decision thresholds. We evaluated agreement and decision-threshold interchangeability between VCTE and 2D-SWE in MASLD with obesity. Methods: In a retrospective cross-sectional agreement study at a tertiary-care center, 317 consecutive adults with MASLD underwent same-day VCTE (FibroScan) and GE LOGIQ E9/E10 2D-SWE. Continuous agreement was assessed by Bland-Altman analysis; the categorical analysis compared binary clinical decision thresholds (VCTE [≥]8.0/[≥]10.0 kPa; 2D-SWE [≥]7.204/[≥]8.060 kPa) using Cohen {kappa} and McNemar tests. Prespecified sensitivity analyses and a post-hoc recalibration were performed; VCTE served as operational reference. Results: VCTE yielded higher values (geometric mean VCTE/2D-SWE ratio, 1.23; ratio limits of agreement, 0.64 2.40; intraclass correlation coefficient, 0.41). At the lower threshold, 71 of 117 VCTE-positive patients (61%) were below the 2D-SWE threshold versus 5 of 200 (2.5%) reclassified upward (agreement 76.0%; {kappa} 0.417; P<.001); the upper threshold was similar (38 of 63, 60%, vs 8 of 254, 3.1%). Discordance persisted across cut-offs; agreement worsened descriptively across BMI strata. Recalibration removed the directional asymmetry but not the discordance (agreement unchanged, 76.0%). Conclusions: In MASLD patients, switching from VCTE to 2D-SWE reclassified decision-threshold status in ~60% of VCTE-positive patients; recalibration removed the directional bias but not the disagreement. Follow-up should use the same elastography modality and, when possible, the same platform.
Aksoy, Y. A.; Lee, S.; Moreno-Bonilla, G.
Show abstract
Background: Cases requiring 13 or more tissue sections in Mohs micrographic surgery (MMS) demand extended operative time, additional resources, and often specialised closure techniques. Pre-operative identification of such cases would improve surgical scheduling, resource allocation, and patient counselling. We aimed to develop and validate a machine learning prediction tool using pre-operative clinical features to identify cases likely to require13 sections. Objectives: To develop and validate machine learning models for predicting which Mohs procedures will require 13 sections, using pre-operative clinical features, and to identify key predictive factors. Methods: We analysed 408 consecutive Mohs procedures with 16 pre-operative clinical variables. Thirty machine learning algorithms were evaluated, including ensemble methods (Stacking, Voting), gradient boosting (XGBoost, LightGBM, CatBoost), neural networks (3-7 layers), support vector machines, and traditional classifiers. Model performance was assessed using 5-fold stratified cross-validation and independent test set evaluation. Feature importance was determined using SHAP (SHapley Additive exPlanations) analysis. Results: The stacking ensemble achieved the highest cross-validation AUC of 0.891 (95% CI: 0.849-0.934) and test AUC of 0.884. Tumour area (cm2), calculated using the ellipse formula to approximate clinical tumour morphology, emerged as the strongest predictor (SHAP importance: 0.141), followed by tumour size dimensions (0.086 and 0.068), aggressive histopathology (0.046), and recurrence status (0.035). Wide neural network architectures (5-layer) outperformed deeper configurations (7-layer). The model demonstrated 70.7% high-confidence predictions with uncertainty <15%. Conclusions: Machine learning models using pre-operative clinical features can accurately predict which Mohs procedures will require 13 or more sections. The stacking ensemble approach provides robust predictions suitable for clinical decision support. External validation in multi-centre cohorts with diverse patient populations and practice patterns is warranted to assess model generalisability.
Pichkar, Y.; Manolakos, S.; Phillips, K. M.; Schabath, M. B.; Chaudhary, A.
Show abstract
Background: Low-dose computed tomography (LDCT) screening reduces lung cancer mortality but is limited by low uptake and associated with high rates of false-positives and indeterminate-nodules. Breath volatile organic compound (VOC) analysis is a non-invasive candidate biomarker approach that could complement LDCT, but prior work has relied on laboratory-based high-resolution mass spectrometry (HRMS), limiting point-of-care deployment. Methods: In this pilot study, breath samples were collected from 40 patients with treatment-naive, pathologically confirmed non-small cell lung cancer (NSCLC) and 25 lung-cancer-screening-eligible healthy controls. Paired samples were analyzed via a compact point-of-care GC-MS platform (CLARION) and a laboratory HRMS reference. Diagnostic classification models were built independently for each platform using elastic net logistic regression with leave-one-out cross-validation, and performance was evaluated by area under the receiver operating characteristic curve (AUC). Results: CLARION identified 103 VOCs across breath specimens, compared to over 900 identified by HRMS. Despite this difference in panel size, CLARION achieved diagnostic performance nearly identical to HRMS for distinguishing NSCLC cases from controls (AUC 0.864 vs. 0.863). Compared to controls, performance statistics were similar for early-stage NSCLC (AUC 0.854 vs. 0.841) and adenocarcinoma (AUC 0.770 vs. 0.787). VOCs of interest include p-cymene, phenol, propylbenzene, tetradecane, {beta}-ocimene, 2,3-dihydro-indole, and 1-methylthio-(Z)-1-propene. Conclusion: A compact, point-of-care breath GC-MS platform achieved diagnostic performance for NSCLC detection comparable to a laboratory HRMS reference despite a substantially smaller detected VOC panel. These findings support continued development of point-of-care breath VOC testing as a non-invasive, field-deployable complement to LDCT-based lung cancer screening.
Posio, R. J. E.; Magpili, K. G.
Show abstract
Breast cancer is the leading cause of cancer-related deaths among women in the Philippines. Over 65% of these cases are diagnosed when they are advanced (Montemayor, 2023). This highlights the need for improved early screening devices. E-HAPLOS, or Electrical Impedance Human-guided Assessment with Pressure for Lump Observation System, is a low-cost glove with sensors designed to improve early detection of suspicious breast lump through touch. It integrates force-sensitive resistors (FSRs) to measure tissue stiffness and Electrical Impedance Spectroscopy (EIS) to analyze conductivity across different frequencies--properties that are closely linked to breast cancer. The prototype uses an ESP32 microcontroller that transmits real-time pressure and impedance data to the website. Tested on gelatin breast models with simulated lump, the FSRs effectively identified lump locations by recording higher mean force values (45.81 kPa vs. 33.57 kPa). This guided approach allowed the combined FSR-EIS system to reach a diagnostic performance with an Area Under the Curve (AUC) above 0.94, a significant improvement over unguided measurement (AUC {approx} 0.78). A two-way ANOVA confirmed a significant difference in diagnostic performance based on the system modality (p < 0.001). Tukeys Honesty Significant Difference (HSD) test showed that the FSR-EIS system was statistically superior to both the unguided EIS (p < 0.001) and FSR-only system (p = 0.041). Results demonstrate the synergistic effect of the integrated system, enabling accurate differentiation of suspicious lumps from normal tissue. The FSR-EIS system of the E-HAPLOS glove shows a great potential for detection of lumps in simulated breasts as a screening tool.
Gonzalez-Moro, I.; Sanchez-Garcia, H.; Medina Cuesta, T.; Rodriguez Lirio, A.; Espin Lopez, M. d. P.; Esquivel Gonzalez, S.; Quintana Ochoa de Alda, E.; de la Pena-Sanz, M.; Marin Cano, L.; Sarasua-Blanco, N.; Ortiz Salinas, P.; Sanfeliu Padulles, A.; Ruiz Adrian, A.; Martinez Isidoro, A.; Aldaiturriaga Otaola, A.; Aramburu Gil, A.; Garcia Gil, A.; Saenz Saenz, A.; Heredia Campos, A.; Fernandez Salado, A.; Ramirez Jarana, A. I.; Tobar Lopez, A. I.; Casarojos Oses, A. J.; Martinez de Maranon Toral, A.; Satiago Hidalgo, A.; Silva Diaz, A.; Basterrechea Miguel, A.; Castanos Lasa, A.; Esteras Vadi
Show abstract
Background: Prospective pregnancy registries and biobanking infrastructures are essential for future translational studies investigating maternal, placental and offspring health. However, circulating nucleic acid analyses are highly sensitive to preanalytical variability, particularly regarding blood-collection tube type and sample processing conditions. We established a prospective pregnancy registry and biobanking workflow at Cruces University Hospital and evaluated the impact of preanalytical variables on circulating cell-free DNA (cfDNA) and cell-free RNA (cfRNA) preservation in maternal plasma collected at delivery. Methods: The Registry of Pregnant Women at Cruces University Hospital was designed as a prospective infrastructure integrating placental sampling, maternal blood collection and ethically controlled future access to maternal and offspring clinical data. Within this framework, peripheral blood samples from 50 women at delivery were simultaneously collected into EDTA, Norgen and Roche tubes. Plasma samples processed within or after 24 hours following collection underwent cfDNA/cfRNA extraction, electrophoretic profiling, fluorometric quantification and RT-qPCR analyses targeting different stress-related genes. Results: By the end of June 2026, 1,127 women had been prospectively recruited into the registry, with 661 plasma samples, 637 serum samples and 858 sets of four placental biopsies collected, processed and stored in the Basque Biobank. In the preanalytical substudy, EDTA tubes yielded higher cfDNA concentrations, likely reflecting reduced cellular preservation and genomic DNA contamination. In contrast, Roche tubes showed superior cfRNA preservation, with higher cfRNA concentrations and more consistent detection of the characteristic 5S rRNA peak compared with EDTA and Norgen tubes. Processing delays beyond 24 hours reduced cfRNA concentration, while associations between circulating transcripts and gestational age were more consistently detectable in preservative-containing tubes. Conclusions: Prospective infrastructures like ours offer strong foundation for large scale, long-term studies in the framework of the Developmental Origins of Health and Disease hypothesis. Technically, Roche tubes provided superior cfRNA preservation and enhanced sensitivity for detecting subtle biological associations, supporting the importance of standardized preanalytical workflows within prospective pregnancy biobanking resource.
Bowers, A.; Elliott, J.; Book, N.; Krishna, S.; Hamburg-Shields, E.
Show abstract
Objective: The purpose of this study was to estimate the negative predictive value (NPV) of screening for group B streptococcus (GBS) colonization in pregnant patients undergoing antepartum hospitalization for GBS colonization status at the time of preterm delivery. Study Design: This prospective, observational cohort study compared GBS colonization status upon initial hospital admission to that at the time of delivery. Pregnant patients at 22 to 35 weeks gestation admitted to the antepartum unit at a tertiary care hospital underwent standard screening for GBS colonization. When preterm labor progressed or iatrogenic preterm delivery was indicated, the GBS colonization test was repeated. Comparison of the sequential test results was performed to determine the NPV of the antepartum screening test for the intrapartum status. Results: 159 eligible patients were enrolled in the study, and 100 completed the study and were included in the analysis. The average gestational age at admission was 30 weeks 1 day (95% confidence interval [CI] 29w4d to 30w6d) and the average duration of pregnancy latency in the study group was 17.5 days (95% CI 15.1 to 19.8). GBS colonization rate at the time of admission was 18% and at the time of delivery was 20%. The NPV of GBS screening at admission was 91.5% (95% CI 83.2 to 96.5%) and the positive predictive value (PPV) was 72.2% (95% CI 46.4 to 90.3%). Conclusion: In a cohort of pregnant patients with preterm pregnancy complications, GBS screening at the time of antepartum hospital admission has an NPV of 91.5% (95% CI 83.2 to 96.5%) for GBS colonization at the time of preterm delivery. This is comparable to the published NPV of routine GBS screening for colonization status at term delivery.
Kurdi, F.; Kurdi, Y.; Kurdi, M.; Pisareva, T. N.; Sukortseva, N. S.; Shiryaev, A. A.; Istranov, A. L.; Reshetov, I. V.
Show abstract
Background. Sentinel lymph node biopsy (SLNB) is an essential component of axillary staging in breast cancer. Fluorescence-guided mapping with indocyanine green (ICG) enables real-time visualization of lymphatic drainage; however, the parameters of fluorescence-signal recording require standardization. Objective. To assess the technical feasibility and clinical applicability of SLNB with ICG under near-infrared (NIR) imaging guidance in patients with breast cancer. Materials and Methods. This prospective single-center study included 30 patients who underwent ICG-guided SLNB between 2023 and 2025. In 24 patients, ICG mapping was combined with technetium-99m radioisotope navigation; in 6 patients, ICG guidance alone was used. The protocol comprised periareolar ICG injection, standardized imaging conditions, and fluorescence-index recording. Results. The protocol was completed in all 30 patients. Fluorescent visualization of the lymphatic pathway and/or the sentinel lymph node (SLN) was achieved in every case, and no ICG-related adverse reactions were recorded. The mean fluorescence index was 213.0 +/- 24.7, the median was 206.0 [192.5-237.2], and the min-max was 180-255. Conclusion. SLNB with ICG under NIR imaging guidance demonstrated technical feasibility in a prospective single-center cohort. Quantitative fluorescence-index recording may serve as a component of standardizing intraoperative fluorescence guidance. Keywords: breast cancer; sentinel lymph node biopsy; indocyanine green; near-infrared imaging; fluorescence lymphography; fluorescence index; axillary staging.
Shrestha, L.; Maharjan, D.; Bista, U.
Show abstract
Objectives: To evaluate the diagnostic accuracy of a publicly available DenseNet-121 convolutional neural network (TorchXRayVision) for triaging chest radiographs of health assessment applicants at a tertiary hospital in Nepal. Design: Prospective, single-centre, shadow-mode diagnostic accuracy validation study. Reported in accordance with the STARD 2015 checklist and STARD-AI/DECIDE-AI guidelines. Setting: Department of Radiology and Imaging, Patan Academy of Health Sciences / Patan Hospital, Lalitpur, Nepal. Participants: 826 consecutive health assessment applicants (foreign employment Pre-Departure Medical Examination and student migration) undergoing chest radiography between 5 June and 20 June 2026. Two cases were excluded due to DICOM technical failure. Index test: DenseNet-121 algorithm (TorchXRayVision library, densenet121-res224-all pretrained weights). A maximum aggregated pathology probability score was derived per radiograph and compared against a post-hoc derived threshold of 0.6258 (selected as the highest threshold achieving the pre-specified >=95% sensitivity criterion). Reference standard: Single-reader-per-case review by one of three radiologists - two board-certified radiodiagnosticians (LS: 276 cases; DM: 275 cases) and one radiology resident (UB: 275 cases) - each blinded to AI output, using a standardised data collection worksheet capturing binary classification (abnormal/normal) and free-text findings. Results: Of 826 radiographs, 41 (4.97%) were classified as abnormal by the reference standard. At the post-hoc derived threshold of 0.6258, the DenseNet-121 algorithm achieved: sensitivity 95.12% (95% CI 83.9-98.7%), specificity 77.2% (95% CI 74.1-80.0%), area under the receiver operating characteristic curve (AUROC) 0.9583 (95% bootstrap CI 0.9225-0.9843), NPV 99.67% (95% Wilson CI 98.8-99.9%), PPV 17.89% (95% Wilson CI 13.4-23.5%), and Cohen's {kappa} 0.237 (95% bootstrap CI 0.174-0.304). Brier score was 0.3621 (null Brier 0.0472) and ECE was 0.564, confirming calibration failure due to score compression (range 0.52-0.72) despite preserved discrimination. Cross-validated results: Ten-fold cross-validation yielded bias-corrected sensitivity 95.12% (95% Wilson CI 83.9-98.7%; optimism 0.00 pp) and specificity 75.80% (95% Wilson CI 72.7-78.7%; optimism +1.40 pp), confirming primary metrics are not materially inflated by circular optimisation. Conclusions: The DenseNet-121 algorithm demonstrated high sensitivity and excellent discrimination for chest radiograph triage in a Nepali health-assessment population, supporting its potential as a rule-out tool (NPV 99.67%). Systematic score compression - preserved discrimination despite calibration shift - is a quantifiable marker of LMIC distributional shift. Prospective local calibration studies are warranted before operational deployment.