Back

Diagnostic Accuracy Of Artificial Intelligence For Analysis Of 1.3 Million Medical Imaging Studies: The Moscow Experiment On Computer Vision Technologies

Vladzymyrskyy, A.; Arzamasov, K.; Gelezhe, P.; Morozov, S.; Ledikhova, N.; Andreychenko, A.; Omelyanskaya, O.; Reshetnikov, R.; Blokhin, I.; Turavilova, E.; Anikina, D.; Kozhikhina, D.; Bondarchuk, D.

2023-08-31 radiology and imaging
10.1101/2023.08.31.23294896 medRxiv
Show abstract

Objectiveto assess the diagnostic accuracy of services based on computer vision technologies at the integration and operation stages in Moscows Unified Radiological Information Service (URIS). Methodsthis is a multicenter diagnostic study of artificial intelligence (AI) services with retrospective and prospective stages. The minimum acceptable criteria levels for the index test were established, justifying the intended clinical application of the investigated index test. The Experiment was based on the infrastructure of the URIS and United Medical Information and Analytical System (UMIAS) of Moscow. Basic functional and diagnostic requirements for the artificial intelligence services and methods for monitoring technological and diagnostic quality were developed. Diagnostic accuracy metrics were calculated and compared. Resultsbased on the results of the retrospective study, we can conclude that AI services have good result reproducibility on local test sets. The highest and at the same time most balanced metrics were obtained for AI services processing CT scans. All AI services demonstrated a pronounced decrease in diagnostic accuracy in the prospective study. The results indicated a need for further refinement of AI services with additional training on the Moscow population datasets. Conclusionsthe diagnostic accuracy and reproducibility of AI services on the reference data are sufficient, however, they are insufficient on the data in routine clinical practice. The AI services that participated in the experiment require a technological improvement, additional training on Moscow population datasets, technical and clinical trials to get a status of a medical device.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.