MetaMuse: A Multi-Agent AI System for Biomedical Metadata Curation and Harmonization
Mittal, E.; Litman, E.; Myers, T.; Agarwal, V.; Gopinath, A.; Kassis, T.
Show abstract
Inconsistent and unstructured metadata in public biomedical repositories, such as the Gene Expression Omnibus (GEO), severely limits data discoverability and research reproducibility. To address this, we introduce MO_SCPLOWETAC_SCPLOWMO_SCPLOWUSEC_SCPLOW, a modular, multi-agent artificial intelligence framework designed to autonomously extract, validate, and standardize unstructured biomedical metadata. Operating through a three-stage architecture utilizing large language model agents, specialized CO_SCPLOWURATORC_SCPLOWAO_SCPLOWGENTSC_SCPLOW contextually extract candidate values for specific target metadata fields. A centralized AO_SCPLOWRBITRATORC_SCPLOWAO_SCPLOWGENTC_SCPLOW enforces cross-field logical consistency to prevent contradictory annotations. Finally, a NO_SCPLOWORMALIZERC_SCPLOWAO_SCPLOWGENTC_SCPLOW leveraging a domain-specific semantic search model (SapBERT) maps these free-text candidates to formal ontological terms. We evaluated MO_SCPLOWETAC_SCPLOWMO_SCPLOWUSEC_SCPLOW on a gold-standard dataset of manually curated GEO samples, achieving over 95% curation accuracy across key target metadata fields, and demonstrated robust scalability on a broader dataset of 400 samples. Notably, MO_SCPLOWETAC_SCPLOWMO_SCPLOWUSEC_SCPLOW avoids data hallucination by defaulting to conservative false negatives when evidence is ambiguous, thereby preserving strict data integrity. By providing a fully auditable and context-aware curation pipeline, MO_SCPLOWETAC_SCPLOWMO_SCPLOWUSEC_SCPLOW offers a scalable solution for enriching public data repositories and accelerating reproducible, data-driven scientific discovery.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Gilda: biomedical entity text normalization with machine-learned disambiguation as a service 94%
- AnnSQL: A Python SQL-based package for fast large-scale single-cell genomics analysis using minimal computational resources 93%
- CoNECo: A Corpus for Named Entity recognition and normalization of protein Complexes 93%
Similar papers in this journal
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 95%
- PEPhub: a database, web interface, and API for editing, sharing, and validating biological sample metadata 95%
- Linking big biomedical datasets to modular analysis with Portable Encapsulated Projects 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.