Assessing Statistical Practices of Existing Artificial Intelligence (AI) Models for Lung Cancer Detection, Prognosis, and Risk Prediction: A Cross-Sectional Meta-Research Study Supplemented by Human and Large Language Model (LLM)-Directed Quality Appraisal
Hou, Y.; Ward, T.; Yang, C.-H.; Jernigan, E.; Caturegli, G.; Boffa, D.; Mukherjee, B.
Show abstract
Artificial intelligence (AI) models with medical images as input data are increasingly proposed to support clinical decisions in lung cancer screening. To assess how these models are developed, evaluated, and reported, and to identify gaps in best statistical practices, we conducted a cross-sectional meta-research study of OpenAlex-indexed studies (January 1, 2023, to June 30, 2025) that developed image-based AI tools to detect lung cancer, predict prognosis, or estimate future risk. Thirty-six studies met our inclusion criteria. Study quality and reporting were appraised using three approaches: subjective ratings from two statisticians and two clinicians, scoring from two AI agents (GPT-5 and Gemini 2.5 Pro), and a guideline-based checklist from the Critical Appraisal and Data Extraction for Systematic Reviews of Prediction Modelling Studies (CHARMS). Convolutional neural networks were used in most of the included studies (69%). Area under the curve was the most frequently reported metric (81%). Our meta-research study also highlights common lapses in these 36 studies, including limited external test set use (39%), insufficient subgroup analyses (28%), and a substantial lack of adherence to established prediction-model reporting guidelines. AI-based quality scoring aligned better with CHARMS-based scores than did human scoring. Spearman correlations with CHARMS were weaker for statisticians/clinicians (p [≤] 0.46) than for the two AI agents (GPT-5 p = 0.66; Gemini 2.5 Pro p = 0.56). Overall, future research should prioritize standardized reporting, use of external test sets, and model performance assessment across subpopulations. Large language models (LLMs) offer a supportive role in providing guideline-driven appraisals to complement human judgment in evaluating AI-based prediction models. 1-2 Sentence DescriptionThis cross-sectional meta-research study synthesizes recent studies that developed artificial intelligence (AI)-driven predictive models using medical images to detect lung cancer, predict prognosis, or estimate future risk, highlighting methodological trends, limitations in model testing and subgroup analyses, and advocating for the need for greater transparency, reliability, quality assessment, and adherence to established reporting guidelines in such studies. Quality assessment of the models carried out by LLMs, human statisticians and clinicians indicates chatbots are more aligned with recommended guidelines than humans.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Surgical Resection, Radiotherapy, And Percutaneous Thermal Ablation for Treatment of Stage 1 Non-Small Cell Lung Cancer: A Systematic Review and Network Meta-Analysis 95%
- Reproducibility and transparency characteristics of oncology research evidence 92%
- GPT for RCTs?: Using AI to measure adherence to reporting guidelines 92%
Similar papers in this journal
- Classification performance bias between training and test sets in a limited mammography dataset 93%
- Accuracy of deep learning based computed tomography diagnostic system of COVID-19: a consecutive sampling external validation cohort study 92%
- ai-corona : Radiologist-Assistant Deep Learning Framework for COVID-19 Diagnosis in Chest CT Scans 92%
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 92%
- Evaluation of an artificial intelligence model for detection of pneumothorax and tension pneumothorax on chest radiograph 92%
- Missing data in the medical record for oncology patients: prevalence and association with outcomes 91%
Similar papers in this journal
- Large-scale validation of the Prediction model Risk Of Bias ASsessment Tool (PROBAST) using a short form: high risk of bias models show poorer discrimination 91%
- The impact of retracted randomised controlled trials on systematic reviews and clinical practice guidelines: a meta-epidemiological study 91%
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 91%
Similar papers in this journal
- Volumetric lung cancer screening reduces unnecessary low-dose computed tomography scans: results from a single-centre prospective trial on 4,119 subjects 94%
- Auto-detection of motion artifacts on CT pulmonary angiograms with a physician-trained AI algorithm 92%
- Deep learning models for poorly differentiated colorectal adenocarcinoma classification in whole slide images using transfer learning 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.