Back

MultiGS: A comprehensive and user-friendly genomic prediction platform Integrating statistical, machine learning, and deep learning models for breeders

You, F.; Zheng, C.; Daniel, J. J. Z.; Li, P.; Taran, B.; Cloutier, S.

2026-01-02 bioinformatics
10.64898/2026.01.02.697306 bioRxiv
Show abstract

Genomic selection (GS) is a core strategy in modern breeding programs, yet the rapid expansion of statistical, machine-learning (ML), and deep-learning (DL) models has made systematic evaluation and practical deployment increasingly challenging. To address these issues, we developed MultiGS, a unified and user-friendly framework that integrates linear, ML, DL, hybrid, and ensemble GS models within a standardized and computationally efficient workflow. MultiGS is implemented through two complementary pipelines: MultiGS-R, a Java/R pipeline implementing 12 statistical and ML models, and MultiGS-P, a Python pipeline integrating 17 models including five linear models, three ML approaches, and nine recently developed DL architectures implemented within the framework. We benchmarked MultiGS using wheat, maize, and flax datasets representing contrasting prediction scenarios. Wheat and maize were evaluated using random training-test splits within the same population, reflecting suitable conditions for assessing model capacity and scalability. Under these scenarios, several DL, hybrid, and ensemble models achieved prediction accuracies comparable to RR-BLUP and consistently exceeded those of GBLUP. In contrast, the flax dataset represented a true across-population prediction scenario with limited training set size and strong population structure. In this challenging context, classical linear models provided stable baselines, while a subset of DL architectures--particularly graph-based models and BLUP-integrated hybrids--demonstrated comparatively improved generalization across populations. Comparisons with previously published DL tools showed that MultiGS models achieved comparable or improved prediction accuracies while requiring lower computational costs, enabling routine retraining and large-scale evaluation. Overall, MultiGS informs, scenario-specific model selection and provides a practical platform for deploying genomic prediction under realistic breeding conditions. The software is freely available on GitHub (https://github.com/AAFC-ORDC-Crop-Bioinfomatics/MultiGS).

Published in Crop Breeding, Genetics and Genomics · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

1
The Plant Genome
57 papers in training set
Top 0.1%
33.2%
2
Theoretical and Applied Genetics
49 papers in training set
Top 0.1%
18.6%
50% of probability mass above
3
Horticulture Research
47 papers in training set
Top 0.2%
4.9%
4
Frontiers in Plant Science
256 papers in training set
Top 2%
3.3%
5
Briefings in Bioinformatics
354 papers in training set
Top 3%
2.7%
6
in silico Plants
27 papers in training set
Top 0.1%
2.7%
7
G3: Genes, Genomes, Genetics
252 papers in training set
Top 2%
2.5%
8
The Plant Phenome Journal
14 papers in training set
Top 0.1%
2.4%
9
G3 Genes|Genomes|Genetics
351 papers in training set
Top 2%
2.4%
10
Molecular Ecology Resources
171 papers in training set
Top 0.9%
2.1%
11
Frontiers in Genetics
230 papers in training set
Top 2%
2.1%
12
Bioinformatics
1204 papers in training set
Top 7%
1.7%
13
GENETICS
483 papers in training set
Top 3%
1.7%
14
BMC Genomics
406 papers in training set
Top 5%
1.5%
15
PLOS ONE
5266 papers in training set
Top 52%
1.4%
16
New Phytologist
346 papers in training set
Top 4%
1.1%
17
Genome Biology
637 papers in training set
Top 7%
1.1%
18
Plant Phenomics
18 papers in training set
Top 0.2%
1.0%
19
GigaScience
212 papers in training set
Top 4%
1.0%
20
The Plant Journal
215 papers in training set
Top 3%
0.8%
21
Bioinformatics Advances
203 papers in training set
Top 4%
0.8%
22
BMC Bioinformatics
457 papers in training set
Top 6%
0.8%
23
Applications in Plant Sciences
23 papers in training set
Top 0.4%
0.8%
24
Scientific Reports
3612 papers in training set
Top 78%
0.6%
25
Plant Physiology
238 papers in training set
Top 3%
0.6%
26
PLOS Computational Biology
1863 papers in training set
Top 22%
0.6%