Back

BioModelsML: Building a FAIR and reproducible collection of machine learning models in life sciences and medicine for easy reuse

Tiwari, D. D.; Hoffmann, N.; Didi, K.; Deshpande, S.; Ghosh, S.; Nguyen, T. V. N.; Raman, K.; Hermjakob, H.; Malik Sheriff, R. S.

2023-05-23 bioinformatics
10.1101/2023.05.22.540599 bioRxiv
Show abstract

Machine learning (ML) models are widely used in life sciences and medicine; however, they are scattered across various platforms and there are several challenges that hinder their accessibility, reproducibility and reuse. In this manuscript, we present the formalisation and pilot implementation of community protocol to enable FAIReR (Findable, Accessible, Interoperable, Reusable, and Reproducible) sharing of ML models. The protocol consists of eight steps, including sharing model training code, dataset information, reproduced figures, model evaluation metrics, trained models, Dockerfiles, model metadata, and FAIR dissemination. Applying these measures we aim to build and share a comprehensive public collection of FAIR ML models in the BioModels repository through incentivized community curation. In a pilot implementation, we curated diverse ML models to demonstrate the feasibility of our approach and we discussed the current challenges. Building a FAIReR collection of ML models will directly enhance the reproducibility and reusability of ML models, minimising the effort needed to reimplement models, maximising the impact on the application and significantly accelerating the advancement in the field of life science and medicine.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 1.0%
22.1%
2
Bioinformatics Advances
203 papers in training set
Top 0.2%
9.8%
3
GigaScience
212 papers in training set
Top 0.2%
9.7%
4
Frontiers in Bioinformatics
49 papers in training set
Top 0.1%
9.0%
50% of probability mass above
5
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.6%
5.2%
6
Database
61 papers in training set
Top 0.2%
4.5%
7
BMC Bioinformatics
457 papers in training set
Top 2%
4.3%
8
Nucleic Acids Research
1281 papers in training set
Top 5%
4.1%
9
PLOS ONE
5266 papers in training set
Top 45%
2.1%
10
Briefings in Bioinformatics
354 papers in training set
Top 4%
2.1%
11
Journal of Molecular Biology
232 papers in training set
Top 2%
1.9%
12
Neuroinformatics
46 papers in training set
Top 0.4%
1.7%
13
F1000Research
88 papers in training set
Top 1%
1.7%
14
PLOS Computational Biology
1863 papers in training set
Top 14%
1.7%
15
Computer Methods and Programs in Biomedicine
28 papers in training set
Top 0.7%
1.1%
16
Frontiers in Genetics
230 papers in training set
Top 4%
1.1%
17
Journal of Open Source Software
25 papers in training set
Top 0.3%
1.1%
18
PeerJ
308 papers in training set
Top 9%
1.0%
19
SoftwareX
15 papers in training set
Top 0.2%
1.0%
20
NAR Genomics and Bioinformatics
242 papers in training set
Top 4%
1.0%
21
BioData Mining
22 papers in training set
Top 0.8%
0.8%
22
International Journal of Molecular Sciences
494 papers in training set
Top 15%
0.8%
23
Frontiers in Neuroinformatics
41 papers in training set
Top 0.6%
0.8%
24
Biology Methods and Protocols
61 papers in training set
Top 2%
0.8%
25
Journal of Cheminformatics
29 papers in training set
Top 0.8%
0.6%