Back

Validation of the QAMAI tool to assess the quality of health information provided by AI.

Vaira, L. A.; Lechien, J. R.; Abbate, V.; Allevi, F.; Audino, G.; Beltramini, G. A.; Bergonzani, M.; Boscolo-Rizzo, P.; Califano, G.; Cammaroto, G.; Chiesa-Estomba, C. M.; Committeri, U.; Crimi, S.; Curran, N. R.; Di Bello, F.; Di Stadio, A.; Frosolini, A.; Gabriele, G.; Gengler, I. M.; Lonardi, F.; Maniaci, A.; Maglitto, F.; Mayo-Yanez, M.; Petrocelli, M.; Pucci, R.; Saibene, A. M.; Saponaro, G.; Tel, A.; Trabalzini, F.; Trecca, E.; Vellone, V.; Salzano, G.; De Riu, G.

2024-01-25 health informatics
10.1101/2024.01.25.24301774 medRxiv
Show abstract

ObjectiveTo propose and validate the Quality Assessment of Medical Artificial Intelligence (QAMAI), a tool specifically designed to assess the quality of health information provided by AI platforms. Study designobservational and valuative study Setting27 surgeons from 25 academic centers worldwide. MethodsThe QAMAI tool has been developed by a panel of experts following guidelines for the development of new questionnaires. A total of 30 responses from ChatGPT4, addressing patient queries, theoretical questions, and clinical head and neck surgery scenarios were assessed. Construct validity, internal consistency, inter-rater and test-retest reliability were assessed to validate the tool. ResultsThe validation was conducted on the basis of 792 assessments for the 30 responses given by ChatGPT4. The results of the exploratory factor analysis revealed a unidimensional structure of the QAMAI with a single factor comprising all the items that explained 51.1% of the variance with factor loadings ranging from 0.449 to 0.856. Overall internal consistency was high (Cronbachs alpha=0.837). The Interclass Correlation Coefficient was 0.983 (95%CI 0.973-0.991; F(29,542)=68.3; p<0.001), indicating excellent reliability. Test-retest reliability analysis revealed a moderate-to-strong correlation with a Pearsons coefficient of 0.876 (95%CI 0.859-0.891; p<0.001) ConclusionsThe QAMAI tool demonstrated significant reliability and validity in assessing the quality of health information provided by AI platforms. Such a tool might become particularly important/useful for physicians as patients increasingly seek medical information on AI platforms.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.