Diagnostic Accuracy Of Artificial Intelligence For Analysis Of 1.3 Million Medical Imaging Studies: The Moscow Experiment On Computer Vision Technologies
Vladzymyrskyy, A.; Arzamasov, K.; Gelezhe, P.; Morozov, S.; Ledikhova, N.; Andreychenko, A.; Omelyanskaya, O.; Reshetnikov, R.; Blokhin, I.; Turavilova, E.; Anikina, D.; Kozhikhina, D.; Bondarchuk, D.
Show abstract
Objectiveto assess the diagnostic accuracy of services based on computer vision technologies at the integration and operation stages in Moscows Unified Radiological Information Service (URIS). Methodsthis is a multicenter diagnostic study of artificial intelligence (AI) services with retrospective and prospective stages. The minimum acceptable criteria levels for the index test were established, justifying the intended clinical application of the investigated index test. The Experiment was based on the infrastructure of the URIS and United Medical Information and Analytical System (UMIAS) of Moscow. Basic functional and diagnostic requirements for the artificial intelligence services and methods for monitoring technological and diagnostic quality were developed. Diagnostic accuracy metrics were calculated and compared. Resultsbased on the results of the retrospective study, we can conclude that AI services have good result reproducibility on local test sets. The highest and at the same time most balanced metrics were obtained for AI services processing CT scans. All AI services demonstrated a pronounced decrease in diagnostic accuracy in the prospective study. The results indicated a need for further refinement of AI services with additional training on the Moscow population datasets. Conclusionsthe diagnostic accuracy and reproducibility of AI services on the reference data are sufficient, however, they are insufficient on the data in routine clinical practice. The AI services that participated in the experiment require a technological improvement, additional training on Moscow population datasets, technical and clinical trials to get a status of a medical device.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- High-Dimensional Multinomial Multiclass Severity Scoring of COVID-19 Pneumonia Using CT Radiomics Features and Machine Learning Algorithms 96%
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 95%
- Identifying relationships between imaging phenotypes and lung cancer-related mutation status: EGFR and KRAS 95%
Similar papers in this journal
- Auto-detection of motion artifacts on CT pulmonary angiograms with a physician-trained AI algorithm 94%
- Passive Microwave Radiometry (MWR) for diagnostics of COVID-19 lung complications in Kyrgyzstan 94%
- Volumetric lung cancer screening reduces unnecessary low-dose computed tomography scans: results from a single-centre prospective trial on 4,119 subjects 92%
Similar papers in this journal
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 95%
- Predicting EGFR mutation status in lung adenocarcinoma presenting as ground-glass opacity: utilizing radiomics model in clinical translation 93%
- Observer agreement and clinical significance of chest CT reporting in patients suspected of COVID-19 93%
Similar papers in this journal
- ai-corona : Radiologist-Assistant Deep Learning Framework for COVID-19 Diagnosis in Chest CT Scans 95%
- REPLICCAR II Study: Data Quality Audit in the Paulista Cardiovascular Surgery Registry 94%
- Navigated ultrasound bronchoscopy with integrated positron emission tomography - A human feasibility study 94%
Similar papers in this journal
- Point-of-care lung ultrasonography for early identification of mild COVID-19: a prospective cohort of outpatients in a Swiss screening center 92%
- Large language model-based information extraction from free-text radiology reports: a scoping review protocol 92%
- Chest X-Ray Has Poor Diagnostic Accuracy and Prognostic Significance in COVID-19: A Propensity Matched Database Study 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.