Autonomous Loop Construction And Supervision For Clinician-Oriented Medical-Ai Research
Chen, X.; Jiang, X.; Shan, C.; Wang, Z.; Li, D.; Zhao, C.
Show abstract
Medical AI models have made a great impact on biomedical research and real-world clinical applications, but conducting interdisciplinary medical AI research remains challenging, requiring close collaboration between clinicians and AI experts. Recent advances in large language models (LLMs) and autonomous code agents present an opportunity for low cost medical AI development, where clinicians can build AI tailored to their own research questions, even without continuous support from dedicated AI experts. However, enabling code agents to autonomously tackle complex multimodal medical AI development tasks requires clinicians to construct and supervise an AI research loop with detailed technical specifics, demanding substantial expertise in AI and computer science that they often lack. To address this challenge, we introduce the Medical AI Research Loop Agent (MARLA), an agentic framework that completely abstracts the construction and supervision of medical AI research loops from clinicians. Given a clinician-defined research intent, MARLA automatically translates high-level research goals into executable hierarchical research loops, decomposes them into verifiable sub-loops, and specifies the models, datasets, tools, and evaluation protocols required for each task. During execution, MARLA coordinates specialized code agents, monitors progress, diagnoses failures, and iteratively refines research strategies based on experimental feedback to drive the research process toward optimal outcomes. We evaluate MARLA on multimodal medical AI tasks that require closed-loop conversion from high-level clinical study objectives to trained and validated AI models. The results demonstrate MARLA's ability to autonomously conduct complex medical AI research while substantially reducing the need for AI expertise.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Interpretable Fine-tuned Large Language Models Facilitate Making Genetic Test Decisions for Rare Diseases 93%
- Zero Shot Health Trajectory Prediction Using Transformer 93%
- Large language models improve transferability of electronic health record-based predictions across countries and coding systems 92%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Generalized Radiograph Representation Learning via Cross-supervision between Images and Free-text Radiology Reports 89%
- COSIME: Cooperative multi-view integration with Scalable and Interpretable Model Explainer 89%
- Estimating Treatment Effects for Time-to-Treatment Antibiotic Stewardship in Sepsis 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.