Large Language Models Can Extract Metadata for Annotation of Human Neuroimaging Publications
Turner, M. D.; Appaji, A.; Ar Rakib, N.; Golnari, P.; Rajasekar, A. K.; Rathnam K V, A.; Sahoo, S. S.; Wang, Y.; Wang, L.; Turner, J. A.
Show abstract
We show that recent (mid-to-late 2024) commercial large language models (LLMs) are capable of good quality metadata extraction and annotation with very little work on the part of investigators for several exemplar real-world annotation tasks in the neuroimaging literature. We investigated the GPT-4o LLM from OpenAI which performed comparably with several groups of specially trained and supervised human annotators. The LLM achieves similar performance to humans, between 0.91 and 0.97 on zero-shot prompts without feedback to the LLM. Reviewing the disagreements between LLM and gold standard human annotations we note that actual LLM errors are comparable to human errors in most cases, and in many cases these disagreements are not errors. Based on the specific types of annotations we tested, with exceptionally reviewed gold-standard correct values, the LLM performance is usable for metadata annotation at scale. We encourage other research groups to develop and make available more specialized "micro-benchmarks," like the ones we provide here, for testing both LLMs, and more complex agent systems annotation performance in real-world metadata annotation tasks.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An open, analysis-ready, and quality controlled resource for pediatric brain white-matter research 94%
- Harmonized diffusion MRI data and white matter measures from the Adolescent Brain Cognitive Development Study 94%
- qMRI-BIDS: an extension to the brain imaging data structure for quantitative magnetic resonance imaging data 93%
Similar papers in this journal
Similar papers in this journal
- Confound modelling in UK Biobank brain imaging 96%
- Highlight Results, Don't Hide Them: Enhance interpretation, reduce biases and improve reproducibility 95%
- Quality control strategies for brain MRI segmentation and parcellation: practical approaches and recommendations - insights from The Maastricht Study 94%
Similar papers in this journal
Similar papers in this journal
- Methods for decoding cortical gradients of functional connectivity 95%
- A Set of FMRI Quality Control Tools in AFNI: Systematic, in-depth and interactive QC with afni_proc.py and more 95%
- BrainQCNet: a Deep Learning attention-based model for the automated detection of artifacts in brain structural MRI scans. 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.