Back

Improving Alphafold2 Performance With A Global Metagenomic & Biological Data Supply Chain

Munsamy, G.; Bohnuud, T.; Lorenz, P.

2024-03-06 genomics
10.1101/2024.03.06.583325 bioRxiv
Show abstract

Scaling laws suggest that more than a trillion species inhabit our planet but only a miniscule and unrepresentative fraction (less than 0.00001%) have been studied or sequenced to date. Deep learning models, including those applied to tasks in the life sciences, depend on the quality and size of training or reference datasets. Given the large knowledge gap we experience when it comes to life on earth, we present a data-centric approach to improving deep learning models in Biology: We built partnerships with nature parks and biodiversity stakeholders across 5 continents covering 50% of global biomes, establishing a global metagenomics and biological data supply chain. With higher protein sequence diversity captured in this dataset compared to existing public data, we apply this data advantage to the protein folding problem by MSA supplementation during inference of AlphaFold2. Our model, BaseFold, exceeds traditional AlphaFold2 performance across targets from the CASP15 and CAMEO, 60% of which show improved pLDDT scores and RMSD values being reduced by up to 80%. On top of this, the improved quality of the predicted structures can yield better docking results. By sharing benefits with the stakeholders this data originates from, we present a way of simultaneously improving deep learning models for biology and incentivising protection of our planets biodiversity.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 2%
16.4%
2
Bioinformatics Advances
203 papers in training set
Top 0.2%
9.8%
3
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.2%
7.3%
4
Artificial Intelligence in the Life Sciences
13 papers in training set
Top 0.1%
7.3%
5
PLOS Computational Biology
1863 papers in training set
Top 5%
6.7%
6
GigaScience
212 papers in training set
Top 0.8%
4.3%
50% of probability mass above
7
Briefings in Bioinformatics
354 papers in training set
Top 2%
4.3%
8
Scientific Reports
3612 papers in training set
Top 26%
4.0%
9
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 0.3%
3.5%
10
PLOS ONE
5266 papers in training set
Top 37%
3.2%
11
BMC Bioinformatics
457 papers in training set
Top 3%
3.2%
12
International Journal of Molecular Sciences
494 papers in training set
Top 4%
3.1%
13
Nature Methods
385 papers in training set
Top 3%
2.8%
14
Journal of Molecular Biology
232 papers in training set
Top 2%
2.0%
15
Journal of Chemical Information and Modeling
238 papers in training set
Top 2%
1.7%
16
Methods
34 papers in training set
Top 0.3%
1.7%
17
International Journal of Biological Macromolecules
76 papers in training set
Top 1%
1.4%
18
NAR Genomics and Bioinformatics
242 papers in training set
Top 4%
1.0%
19
Frontiers in Genetics
230 papers in training set
Top 5%
1.0%
20
Communications Biology
993 papers in training set
Top 30%
0.8%
21
Nucleic Acids Research
1281 papers in training set
Top 15%
0.6%
22
Nature Machine Intelligence
70 papers in training set
Top 3%
0.6%