Back

MMMAS: A Mendelian Mismatch Matrix Analysis System for Deterministic Pre-Screening of Germplasm Collections

Chen, Q.; chen, L.; Fan, J.; Yang, X.; Zhang, J.; Du, W.; Du, W.; Liu, Z.; Hu, H.

2026-08-06 genetics
10.64898/2026.08.02.742285 bioRxiv
Show abstract

Germplasm collections are expanding, yet erroneous pedigree records, duplicate accessions, and undocumented kinship remain widespread. We present MMMAS (Mendelian Mismatch Matrix Analysis System) v1.0.0, an open-source tool that converts pairwise Mendelian mismatch numbers (Mmn) into an N x N mismatch matrix. From this matrix, MMMAS derives five population-scale diagnostics: the Mendelian Minimum Mismatch Number (Mmin), the Mendelian Average Mismatch Number (Mavg), the Mendelian Zero-mismatch Partner Number (Mzmp), the Mendelian Mismatch Mode Duplication Index (Mmmd), and the Mendelian Exhaustive Stratification (MES) grading system, with the core computation requiring neither allele frequency estimates nor assumptions of genetic models. When validated on 1,085 apple and 383 sweet cherry accessions from the German Fruit Genebank, MMMAS reproduced published CERVUS assignments at 97.31% (apple)[~]100% (sweet cherry) recall, detected one literature-confirmed duplicate genotype via the Mmmd index, and identified documented breeding hub parents such as Cox Orange. Marker-reduction analysis on the apple dataset showed that nine simple sequence repeat (SSR) loci were sufficient to maintain a stable MES grading structure, whereas resolving highly distinct wild germplasm required no fewer than 13 loci. End-to-end analysis of a simulated 10,000-accession panel (50 million pairwise comparisons) completed in under six minutes on a workstation, with runtime scaling near-linearly with the number of pairwise comparisons (O(N2L)). MMMAS is released under the MIT license, featuring a bilingual graphical interface, standard CSV input, and complete documentation. It provides a deterministic pre-screening layer for stratifying genetic distinctness in germplasm collections.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.