Validation of the QAMAI tool to assess the quality of health information provided by AI.
Vaira, L. A.; Lechien, J. R.; Abbate, V.; Allevi, F.; Audino, G.; Beltramini, G. A.; Bergonzani, M.; Boscolo-Rizzo, P.; Califano, G.; Cammaroto, G.; Chiesa-Estomba, C. M.; Committeri, U.; Crimi, S.; Curran, N. R.; Di Bello, F.; Di Stadio, A.; Frosolini, A.; Gabriele, G.; Gengler, I. M.; Lonardi, F.; Maniaci, A.; Maglitto, F.; Mayo-Yanez, M.; Petrocelli, M.; Pucci, R.; Saibene, A. M.; Saponaro, G.; Tel, A.; Trabalzini, F.; Trecca, E.; Vellone, V.; Salzano, G.; De Riu, G.
Show abstract
ObjectiveTo propose and validate the Quality Assessment of Medical Artificial Intelligence (QAMAI), a tool specifically designed to assess the quality of health information provided by AI platforms. Study designobservational and valuative study Setting27 surgeons from 25 academic centers worldwide. MethodsThe QAMAI tool has been developed by a panel of experts following guidelines for the development of new questionnaires. A total of 30 responses from ChatGPT4, addressing patient queries, theoretical questions, and clinical head and neck surgery scenarios were assessed. Construct validity, internal consistency, inter-rater and test-retest reliability were assessed to validate the tool. ResultsThe validation was conducted on the basis of 792 assessments for the 30 responses given by ChatGPT4. The results of the exploratory factor analysis revealed a unidimensional structure of the QAMAI with a single factor comprising all the items that explained 51.1% of the variance with factor loadings ranging from 0.449 to 0.856. Overall internal consistency was high (Cronbachs alpha=0.837). The Interclass Correlation Coefficient was 0.983 (95%CI 0.973-0.991; F(29,542)=68.3; p<0.001), indicating excellent reliability. Test-retest reliability analysis revealed a moderate-to-strong correlation with a Pearsons coefficient of 0.876 (95%CI 0.859-0.891; p<0.001) ConclusionsThe QAMAI tool demonstrated significant reliability and validity in assessing the quality of health information provided by AI platforms. Such a tool might become particularly important/useful for physicians as patients increasingly seek medical information on AI platforms.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 94%
- The NASSS (Non-Adoption, Abandonment, Scale-Up, Spread and Sustainability) framework use over time: A scoping review 94%
- Defining Destigmatizing Design Guidelines for Use in Sexual Health-Related Digital Technologies: A Delphi Study 94%
Similar papers in this journal
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 95%
- Health indicators as a measure of individual health status: public perspectives 94%
- Design and implementation of a system for automated monitoring of adherence to evidenced-based clinical guideline recommendations 94%
Similar papers in this journal
- REPLICCAR II Study: Data Quality Audit in the Paulista Cardiovascular Surgery Registry 94%
- Evaluating user experience with immersive technology in simulation-based education: a modified Delphi study with qualitative analysis 94%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 94%
Similar papers in this journal
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 94%
- Measures of socioeconomic advantage are not independent predictors of support for healthcare AI: subgroup analysis of a national Australian survey 94%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 92%
Similar papers in this journal
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 95%
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 95%
- Assessment of Accuracy and Safety of LabTest Checker (LTC-AI) 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.