Anatomical Accuracy of Generative AI for Congenital Heart Disease Illustrations: Gemini NanoBanana Versus ChatGPT Models in a Blinded Comparative Study
Alhuzaimi, A.; Alkanhal, A.; Alruwaili, A. R. S.; Alharbi, N. S.; Alfares, F.; Aldekhyyel, R. N.; Binkheder, S.; Temsah, A.; Aljamaan, F.; Shahzad, M.; Albriek, A. Z.; Alanazi, F. I.; Alhindi, D. A.; Al-khatib, S. M.; Darweesh, A. A.; Altamimi, I.; Jamal, A.; Saad, K.; Alhasan, K.; Al-Eyadhy, A.; Malki, K. H.; Temsah, M.-H.
Show abstract
BackgroundGenerative artificial intelligence (AI) systems are increasingly used to produce medical illustrations for education; however, their anatomical accuracy in complex domains such as congenital heart disease (CHD) remains insufficiently validated. MethodsIn an assessor-blinded comparative study, we evaluated AI-generated CHD illustrations from two contemporary text-to-image platforms (ChatGPT-5/ChatGPT-Images and Gemini NanoBanana) against human-modified educational images. Twenty different CHD types were included, yielding 147 images that were assessed by 20 physicians (10 CHD experts and 10 non-specialists). Images were rated across four domains: anatomical accuracy, label usefulness, visual attractiveness, and suitability for medical education (total score range, 4-12). ResultsAmong 2,940 total image evaluations, the human-modified images demonstrated the highest anatomical accuracy (48.3% rated accurate), followed by NanoBanana (22.7%), while ChatGPT-generated images were predominantly rated as fabricated or incorrect (86.3% for ChatGPT-5 and 85.2% for ChatGPT-Images; p<0.001). Educational usability "as is" was highest for the human-modified images (37.9%) compared with NanoBanana (13.1%) and ChatGPT platforms ([≤]2.1%; p<0.001). Median overall quality scores were 8 for the human-modified CHD images and NanoBanana, versus 4 for both ChatGPT systems (p<0.001). In multivariable analysis, NanoBanana images were the closest to the human-modified images in quality (95% CI, 0.91-0.98), while ChatGPT-Images (95% CI, 0.58-0.63) and ChatGPT-5 (95% CI, 0.55-0.59) showed marked quality reductions. ConclusionsThe current generative AI systems produced visually compelling but frequently anatomically inaccurate CHD illustrations, falling substantially short of the current educational standards. Model choice strongly influences performance, with Gemini NanoBanana outperforming ChatGPT-based systems yet remaining inferior to expert-designed human-modified images. AI-generated cardiac imagery should be used only within expert-reviewed educational workflows rather than as independent instructional resources.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 95%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 93%
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 93%
Similar papers in this journal
- ChatGPT Provides Inconsistent Risk-Stratification of Patients With Atraumatic Chest Pain 92%
- {-}CardiOvascular examination in awake Orangutans (Pongo pygmaeus pygmaeus): Low-stress Echocardiography including Speckle Tracking imaging (the COOLEST method) 92%
- Evaluating user experience with immersive technology in simulation-based education: a modified Delphi study with qualitative analysis 92%
Similar papers in this journal
- Evaluating Text-to-Image Generated Photorealistic Images of Human Anatomy 92%
- “This is a quiz” Premise Input: A Key to Unlocking Higher Diagnostic Accuracy in Large Language Models 91%
- Referral for Cardiac Amyloidosis in Patients who underwent Transcatheter Aortic Valve Replacement: Result of Quality Outcome Project 90%
Similar papers in this journal
- GenECG: A synthetic image-based ECG dataset to augment artificial intelligence-enhanced algorithm development 96%
- Measures of socioeconomic advantage are not independent predictors of support for healthcare AI: subgroup analysis of a national Australian survey 90%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 89%
Similar papers in this journal
- Aortic annuloplasty FSI digital twin of 3D-printed phantoms with 4D-flow MRI comparison 92%
- Deployment of a digital twin using the coupled momentum method for fluid-structure interaction: a case study for aortic aneurysm 91%
- Detecting Heart Failure using novel bio-signals and a Knowledge Enhanced Neural Network 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.