Autonomous Agents for Auditable Cardiovascular Artificial Intelligence Development
Dhingra, L. S.; Batinica, B.; Choi, R. B.; Croon, P. M.; Oikonomou, E. K.; Khera, R.
Show abstract
Clinical artificial intelligence (AI) models are usually reported as finished artifacts, but each model reflects a limited human search across a much larger space of architectures, inputs, losses, optimizers, and training recipes. We tested whether autonomous code-writing agents could perform a controlled model-development experiment: proposing and evaluating code changes, and seeking performance gains without new data or human-guided edits. We built two such agents: an Iteration Agent that searches sequentially, keeping the best variant at each step, and an Evolution Agent that searches for variations in parallel using multiple large language models and prioritizes high-performing lineages across generations. In two architecturally distinct AI-enhanced electrocardiography (AI-ECG) models for structural heart disease, agent-optimized variants improved rank discrimination across held-out, external, and cross-institution evaluations, with area under the receiver operating characteristic curve gains of +0.006 to +0.039 (paired p < 0.05). At a fixed 90% sensitivity, specificity rose by up to 7.1 percentage points and positive predictive value by up to 4.8 percentage points. The selected code changes were substantive, spanning architecture, representation, and training recipe variations. These findings position autonomous agents as an auditable layer for clinical AI model improvement, provided that candidate selection, external validation, and post-update governance are explicit. We release these agents as an open, reusable toolkit.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 93%
- Zero Shot Health Trajectory Prediction Using Transformer 93%
- Cohort Design and Natural Language Processing to Reduce Bias in Electronic Health Records Research: The Community Care Cohort Project 93%
Similar papers in this journal
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 94%
- EHR Foundation Models Improve Robustness in the Presence of Temporal Distribution Shift 94%
- Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones 93%
Similar papers in this journal
Similar papers in this journal
- Deep Learning Prediction of Biomarkers from Echocardiogram Videos 93%
- Predicting the functional effects of voltage-gated potassium channel missense variants with multi-task learning 93%
- Transformer-based deep learning model for the diagnosis of suspected lung cancer in primary care based on electronic health record data 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.