Back

Artificial Intelligence Powered Research Automation (AIPRA) Versus Human Expert: A Two-Arm Ophthalmology Comparative Study

Musleh, A.; Alwisi, N.; Abu Serhan, H.; Toubasi, A.; Malkawi, L.; Alryalat, S. A.

2025-10-29 health informatics
10.1101/2025.10.27.25338904 medRxiv
Show abstract

PurposeTo compare the quality and efficiency of an AI-powered research automation (AIPRA) workflow with a conventional human-led workflow for producing a full systematic review manuscript on the same question. MethodsTwo independent pipelines (human-led vs. AIPRA) each generated a complete manuscript addressing "What is the role of large language models in glaucoma diagnosis?". No protocols or templates were shared. Three blinded domain experts rated five domains on 5-point Likert scales. The primary endpoint was the overall quality of each workflow from query to final manuscript. ResultsMean total scores: human 74.7%, AIPRA 65.3%. The mean difference (AIPRA - Human) was -9.3% (95% CI, -18.8% to 0.0%), meeting the pre-specified non-inferiority criterion. Domain means were identical for query development (66.7% each); the human-led pipeline scored higher in screening, field selection, full-text extraction, and manuscript writing. AIPRA completed the workflow in approximately 2 hours versus about 1 month for the human pipeline (375x faster). ConclusionAIPRA was non-inferior to human experts on overall quality while drastically reducing time to completion. Appropriate human oversight remains important, especially for screening and extraction tasks.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.