Standardised evaluation and monitoring of site-specific AI performance with physical CT phantoms
Genske, U.; Laudani, A.; Yan, L.; Peng, Y.; Boening, G.; Ulas, S. T.; Wagner, M. P.; Diekhoff, T.; Hamm, B.; Jahnke, P.
Show abstract
Artificial intelligence (AI) applications in computed tomography (CT) imaging require objective and continuous testing, yet standardised methods for this purpose have not been established. Here, we present a framework using physical phantoms for standardised testing and monitoring of AI, demonstrated in liver lesion detection. We begin by designing phantoms tailored to the anatomical input domain expected by AI algorithms, and then systematically assess how AI performance is affected by variations in scanner technology and operation across two clinical CT systems. Next, we perform longitudinal monitoring, yielding consistent results over fifteen months on both systems. Finally, we validate clinical relevance by demonstrating that AI models trained on phantom data generalize effectively to patients and exhibit no evidence of phantom-specific adaptation. Our findings show that anatomically realistic phantoms enable standardised, site-specific testing and monitoring of AI, providing a proactive method for local and cross-institutional quality assurance.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Dual Adversarial Deconfounding Autoencoder for joint batch-effects removal from multi-center and multi-scanner radiomics data 94%
- Improving Rectal Tumor Segmentation with Anomaly Fusion Derived from Anatomical Inpainting: A Multicenter Study 93%
- An AI-based segmentation and analysis pipeline for high-field MR monitoring of cerebral organoids 92%
Similar papers in this journal
- Reproducible spectral CT thermometry with liver-mimicking phantoms for image-guided thermal ablation 95%
- Patient-derived PixelPrint phantoms for evaluating clinical imaging performance of a deep learning CT reconstruction algorithm 93%
- Detection of Transient Neurotransmitter Response Using Personalized Neural Networks 93%
Similar papers in this journal
- First-generation clinical dual-source photon-counting CT: ultra-low dose quantitative spectral imaging 92%
- From Community Acquired Pneumonia to COVID-19: A Deep Learning Based Method for Quantitative Analysis of COVID-19 on thick-section CT Scans 91%
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.