Development and automated deployment of a specialised machine learning schema within a collaborative research centre: an explorative approach using large language models
Kaier, K.; Benadi, G.; Nolde, S.; Tagle Ludwig, C.; Giuliani, C.; Engel, F.; Watter, M.; Binder, H.
Show abstract
Achieving interoperability in machine learning (ML) workflows remains a significant challenge due to the heterogeneity of data types, algorithms, and application domains, as well as the lack of standardized metadata. In this study, we present the development of a specialized ML metadata schema within the context of the Small Data Initiative, a Collaborative Research Center characterized by diverse scientific approaches. We employed an interdisciplinary process combining expert input, iterative refinement, and schema validation using large language models (LLMs). A two-step LLM-based annotation methodology was applied to 14 representative scientific publications, using six different LLMs to identify both predefined (step 1) and additional ML-related metadata elements (step 2). Manual validation through face-to-face interviews with the main authors to the publications confirmed high precision rates of 70%-85% in the initial step and 86%-98% in the second step, with notable performance variation across models. This approach enabled both the identification of schema inconsistencies and the integration of previously overlooked concepts, leading to the refinement of the metadata schema. The process supports an "AI-by-design" paradigm, ensuring that metadata schemas and annotation workflows are optimized from the outset for downstream AI/ML applications. Our findings also highlight the value of LLM benchmarking in selecting suitable models for domain-specific tasks. Overall, the proposed methodology enhances metadata quality, fosters reproducibility, and contributes to making research data more AI-ready.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- COVID-19-related research data availability and quality according to the FAIR principles: A meta-research study 94%
- Academic Tracker: Software for Tracking and Reporting Publications Associated with Authors and Grants 94%
- The Rise of Open Data Practices Among Bioscientists at the University of Edinburgh 94%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- A Novel Question-Answering Framework for Automated Abstract Screening Using Large Language Models 95%
- Automating literature screening and curation with applications to computational neuroscience 95%
- Annotation-preserving machine translation of English corpora to validate Dutch clinical concept extraction tools 94%
Similar papers in this journal
- Quantitative monitoring of nucleotide sequence data from genetic resources in context of their citation in the scientific literature 94%
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 92%
- FAIR Data Station for Lightweight Metadata Management & Validation of Omics Studies 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.