Comparing human and AI performance in medical machine learning: An open-source Python library for the statistical analysis of reader study data
McKinney, S. M.
Show abstract
In seeking to understand the potential effects of artificial intelligence (AI) on the practice of diagnostic medicine, many investigations involve collecting interpretations from several human experts on a common set of cases. In an effort to standardize the process of analyzing the data emerging from such studies, we have released an open-source Python library to perform applicable statistical procedures. The software implements the industry-standard Obuchowski-Rockette-Hillis (ORH) method for multi-reader multi-case (MRMC) studies. The tools can be used to compare a standalone algorithm against a panel of readers, or compare readers operating in two modalities (for example, with and without algorithmic assistance). The software supports both nonequivalence and noninferiority tests. Functions are also provided to simulate reader and model scores, useful for Monte Carlo power analysis. The code is publicly available in our Gitub repository at https://github.com/Google-Health/google-health/tree/master/analysis.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Implementation and prospective real-time evaluation of a generalized system for in-clinic deployment and validation of machine learning models in radiology 95%
- Classification of Hyper-scale Multimodal Imaging Datasets 93%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 93%
Similar papers in this journal
- Analysis of Eligibility Criteria Clusters Based on Large Language Models for Clinical Trial Design 92%
- ENRICHing Medical Imaging Training Sets Enables More Efficient Machine Learning 92%
- Quantification of abdominal fat from computed tomography using deep learning and its association with electronic health records in an academic biobank 92%
Similar papers in this journal
- Equipping Computational Pathology Systems with Artifact Processing Pipelines: A Showcase for Computation and Performance Trade-offs 94%
- Towards a Clinically-based Common Coordinate Framework for the Human Gut Cell Atlas - The Gut Models 91%
- Addressing Label Noise for Electronic Health Records: Insights from Computer Vision for Tabular Data 91%
Similar papers in this journal
- pyKNEEr: An image analysis workflow for open and reproducible research on femoral knee cartilage 93%
- ai-corona : Radiologist-Assistant Deep Learning Framework for COVID-19 Diagnosis in Chest CT Scans 93%
- Enhancing Semantic Segmentation in Chest X-Ray Images through Image Preprocessing: ps-KDE for Pixel-wise Substitution by Kernel Density Estimation 93%
Similar papers in this journal
- Advancing data science in drug development through an innovative computational framework for data sharing and statistical analysis 92%
- Prediction-powered Inference for Clinical Trials 91%
- External control arm analysis: an evaluation of propensity score approaches, G-computation, and doubly debiased machine learning 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.