Beyond the Black Box: Avenues to Transparency in Regulating Radiological AI/ML-enabled SaMD via the FDA 510(k) Pathway
Youssef, A. T.; Fronk, D.; Grimes, J. N.; Cheuy, L.; Larson, D. B.
Show abstract
BackgroundThe majority of AI/ML-enabled software as a medical device (SaMD) has been cleared through the FDA 510(k) pathway, but with limited transparency on algorithm development details. Because algorithm quality depends on the quality of the training data and algorithmic input, this study aimed to assess the availability of algorithm development details in the 510(k) summaries of AI/ML-enabled SaMD. Then, clinical and/or technical equivalence between predicate generations was assessed by mapping the predicate lineages of all cleared computer-assisted detection (CAD) devices, to ensure equivalence in diagnostic function. MethodsThe FDAs public database was searched for CAD devices cleared through the 510(k) pathway. Details on algorithmic input, including annotation instructions and definition of ground truth, were extracted from summary statements, product webpages, and relevant publications. These findings were cross-referenced with the American College of Radiology-Data Science Institute AI Central database. Predicate lineages were also manually mapped through product numbers included within the 510(k) summaries. ResultsIn total, 98 CAD devices had been cleared at the time of this study, with the majority being computer-assisted triage (CADt) devices (67/98). Notably, none of the cleared CAD devices provided image annotation instructions in their summaries, and only one provided access to its training data. Similarly, more than half of the devices did not disclose how the ground truth was defined. Only 13 CAD devices were reported in peer-reviewed publications, and only two were evaluated in prospective studies. Significant deviations in clinical function were seen between cleared devices and their claimed predicate. ConclusionThe lack of imaging annotation instructions and signicant mismatches in clinical function between predicate generations raise concerns about whether substantial equivalence in the 510(k) pathway truly equates to equivalent diagnostic function. Avenues for greater transparency are needed to enable independent evaluations of safety and performance and promote trust in AI/ML-enabled devices.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Classification performance bias between training and test sets in a limited mammography dataset 93%
- Navigated ultrasound bronchoscopy with integrated positron emission tomography - A human feasibility study 92%
- Enhancing Semantic Segmentation in Chest X-Ray Images through Image Preprocessing: ps-KDE for Pixel-wise Substitution by Kernel Density Estimation 91%
Similar papers in this journal
- Regulatory-approved Deep Learning/Machine Learning-Based Medical Devices in Japan as of 2020: A Systematic Review 95%
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 93%
- Implementation and prospective real-time evaluation of a generalized system for in-clinic deployment and validation of machine learning models in radiology 93%
Similar papers in this journal
- Evaluation of an artificial intelligence model for detection of pneumothorax and tension pneumothorax on chest radiograph 92%
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 89%
- Characterizing Potential Conflicts of Interest Among UpToDate and DynaMed Content Contributors 89%
Similar papers in this journal
Similar papers in this journal
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 92%
- MyoVision-US: an Artificial Intelligence-Powered Software for Automated Analysis of Skeletal Muscle Ultrasonography 91%
- Toward Understanding COVID-19 Pneumonia: A Deep-learning-based Approach for Severity Analysis and Monitoring the Disease 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.