Advancing FAIR Data Management through AI-Assisted Curation of Morphological Data Matrices
Jariwala, S.; Long-Fox, B. L.; Berardini, T. Z.
Show abstract
Curation of biological and paleontological datasets is a labor-intensive process that requires standardization and validation to ensure data integrity. In particular, manual curation of datasets is prone to human errors such as typographical errors, inconsistent formatting, and incomplete metadata, which hinder reproducibility and compliance with Findability, Accessibility, Interoperability, and Reusability (FAIR) principles. Artificial Intelligence (AI) offers a transformative solution for enhancing research efficiency by automating data validation, improving accuracy, and streamlining curation workflows. This study explores the integration of an AI-assisted curation tool developed for MorphoBank, an open access repository established to enhance standardization and usability of morphological character datasets. Specifically, this work presents an AI tool designed to extract, structure, and standardize morphological character data from published literature into the NEXUS file format, a widely used format for phylogenetic analyses. This tool leverages machine learning techniques, including Large Language Models (LLMs), to automate the extraction of character names and states from text in various formats, reducing manual data entry errors and improving data completeness. The system enables efficient conversion of matrix-only files into complete, machine- and human-readable datasets that include key character metadata. By automating these tasks, the tool significantly accelerates dataset curation while improving accuracy and standardization. This approach increases the FAIRness of the data and offers a scalable framework for extending AI-assisted curation to standardize other biological datasets. Our findings demonstrate the value of AI in scientific data curation and advancing data reuse in paleontology, systematics, and evolutionary biology.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 94%
- Sequence Compression Benchmark (SCB) database - a comprehensive evaluation of reference-free compressors for FASTA-formatted sequences 94%
- Quantitative monitoring of nucleotide sequence data from genetic resources in context of their citation in the scientific literature 94%
Similar papers in this journal
- Academic Tracker: Software for Tracking and Reporting Publications Associated with Authors and Grants 95%
- Geographic Name Resolution Service: A tool for the standardization and indexing of world political division names, with applications to species distribution modeling 95%
- Datavzrd: Rapid programming- and maintenance-free interactive visualization and communication of tabular data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.