Back

An Agentic, No Code Artificial Intelligence Workflow for Developing and Externally Validating a Thyroid Nodule Ultrasound Malignancy Classifier

Thomas, J.; Pozdeyev, N.

2026-06-26 endocrinology
10.64898/2026.06.23.26356395 medRxiv
Show abstract

Convolutional neural networks (CNNs) can classify thyroid nodules on ultrasound, yet published models are seldom available for independent testing, require machine learning expertise to develop and deploy, and are validated mostly on papillary thyroid carcinoma. Objective. To test whether an autonomous (agentic), no code artificial intelligence (AI) agent can develop a calibrated thyroid-nodule malignancy classifier, and to validate it internally and on an external cohort spanning multiple cancer histologies. Methods. This is a retrospective, computational diagnostic study with prespecified endpoints. A no code agent (Hugging Face ML Intern) autonomously reviewed data, selected and trained the model and calibrated probabilities, using the open source TN5000 dataset (3500 training, 500 validation, and 1000 test images). The trained ResNet 18 model was externally validated on 232 nodules from the University of Colorado, including follicular, medullary, oncocytic, and follicular variant of papillary carcinomas. Results. On the internal test set, an agentic AI model achieved AUROC 0.94 (95% CI, 0.920 - 0.953), sensitivity 0.90, and specificity 0.80. On external validation, agentic AI model achieved an AUROC of 0.90 (95% CI, 0.850 - 0.936), sensitivity of 0.92, and specificity of 0.68, negative predictive value of 0.96, and positive predictive value of 0.52, exceeding the performance of a previously published classifier on the same cohort (AUROC of 0.83). Conclusions. An agentic, no code AI workflow produced a calibrated, externally validated thyroid nodule classifier, supporting accessible, reproducible, and independently testable medical AI development. Prospective validation and local recalibration are required before clinical use.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Communications Medicine
113 papers in training set
Top 0.1%
15.4%
2
Journal of Pathology Informatics
15 papers in training set
Top 0.1%
12.9%
3
npj Digital Medicine
118 papers in training set
Top 0.6%
11.2%
4
Biology Methods and Protocols
61 papers in training set
Top 0.1%
7.4%
5
Scientific Reports
3612 papers in training set
Top 12%
6.4%
50% of probability mass above
6
PLOS ONE
5266 papers in training set
Top 36%
3.5%
7
PLOS Global Public Health
344 papers in training set
Top 4%
3.3%
8
JMIR Public Health and Surveillance
45 papers in training set
Top 0.3%
2.7%
9
JMIR Medical Informatics
18 papers in training set
Top 0.3%
2.7%
10
PLOS Digital Health
106 papers in training set
Top 2%
2.5%
11
Computational and Structural Biotechnology Journal
242 papers in training set
Top 3%
1.9%
12
Nature Communications
5641 papers in training set
Top 44%
1.8%
13
JAMIA Open
42 papers in training set
Top 0.9%
1.7%
14
eBioMedicine
183 papers in training set
Top 3%
1.5%
15
Diagnostics
50 papers in training set
Top 2%
1.1%
16
BMC Medicine
176 papers in training set
Top 3%
1.1%
17
BMC Medical Research Methodology
47 papers in training set
Top 1%
1.1%
18
Frontiers in Medicine
120 papers in training set
Top 3%
1.1%
19
JCO Clinical Cancer Informatics
22 papers in training set
Top 0.6%
1.0%
20
Journal for ImmunoTherapy of Cancer
75 papers in training set
Top 2%
0.9%
21
Schizophrenia
21 papers in training set
Top 0.4%
0.9%
22
Journal of the American Medical Informatics Association
71 papers in training set
Top 2%
0.9%
23
PeerJ
308 papers in training set
Top 12%
0.6%
24
Journal of Clinical Medicine
97 papers in training set
Top 5%
0.6%
25
Frontiers in Physiology
106 papers in training set
Top 3%
0.6%
26
BMJ Open
601 papers in training set
Top 13%
0.6%
27
The Journal of Pathology
26 papers in training set
Top 0.9%
0.6%
28
Cancer Research Communications
51 papers in training set
Top 2%
0.6%