Back

GENE-FAM: An automated pipeline for mining gene families and its application to MADS-box genes in Cannabis sativa

Ryan, L.; Trubanova, N.; Pender, G.; Melzer, R.; Hughes, G. M.; Schilling, S.

2026-06-15 genomics
10.64898/2026.06.10.731441 bioRxiv
Show abstract

Understanding how gene families evolve can offer great insight into adaptation at the phenotypic and ecological levels. This is particularly true in plants, where transcription factor gene families are often targeted for breeding programs to improve the agronomic traits of economically important crops. While recent advances in next generation sequencing have accelerated the wealth of genomics data, there remains a lack of accessible and reproducible genome mining pipelines tailored for gene family characterisation. Here, we address this gap by developing GENE-FAM, an automated, scalable and open-source pipeline designed to mine and predict gene families based on conserved domains and motifs. To illustrate its application, we apply GENE-FAM to annotate MADS-box transcription factor genes across multiple Cannabis sativa genomes. A comprehensive set of MADS-box genes was identified across three C. sativa cultivars, including both previously annotated and newly predicted genes. Through phylogenetic analyses, we confirm that all type II MADS-box gene subfamilies represented in flowering plants are present in C. sativa. Comparing our annotations with those of Arabidopsis thaliana and Solanum lycopersicum revealed that while most MADS type II families are highly conserved, SEPALLATA-like genes have undergone diversification in C. sativa. Together, these results demonstrate the application of GENE-FAM for genome-wide identification and characterisation of gene families in non-model species, revealing novel insights into MADS-box gene family evolution in C. sativa.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

1
Plant Communications
36 papers in training set
Top 0.1%
12.7%
2
The Plant Genome
57 papers in training set
Top 0.1%
9.5%
3
Plant Biotechnology Journal
64 papers in training set
Top 0.2%
6.5%
4
BMC Genomics
406 papers in training set
Top 0.9%
6.1%
5
Scientific Reports
3612 papers in training set
Top 18%
5.3%
6
GigaScience
212 papers in training set
Top 0.7%
4.7%
7
Plant Direct
95 papers in training set
Top 0.7%
4.2%
8
New Phytologist
346 papers in training set
Top 2%
3.9%
50% of probability mass above
9
G3: Genes|Genomes|Genetics
35 papers in training set
Top 0.1%
3.9%
10
Frontiers in Plant Science
256 papers in training set
Top 2%
3.1%
11
The Plant Journal
215 papers in training set
Top 2%
3.1%
12
Genome Biology
637 papers in training set
Top 4%
3.1%
13
Applications in Plant Sciences
23 papers in training set
Top 0.2%
2.3%
14
The Plant Cell
161 papers in training set
Top 1%
2.3%
15
PLOS ONE
5266 papers in training set
Top 46%
2.1%
16
Nucleic Acids Research
1281 papers in training set
Top 9%
1.7%
17
G3: Genes, Genomes, Genetics
252 papers in training set
Top 3%
1.4%
18
Computational and Structural Biotechnology Journal
242 papers in training set
Top 4%
1.3%
19
PLOS Computational Biology
1863 papers in training set
Top 18%
1.1%
20
PLANTS, PEOPLE, PLANET
27 papers in training set
Top 0.5%
1.1%
21
Plant Physiology
238 papers in training set
Top 3%
1.1%
22
Horticulture Research
47 papers in training set
Top 0.7%
1.1%
23
Plant Molecular Biology
20 papers in training set
Top 0.7%
1.0%
24
NAR Genomics and Bioinformatics
242 papers in training set
Top 4%
1.0%
25
Scientific Data
209 papers in training set
Top 2%
1.0%
26
BMC Bioinformatics
457 papers in training set
Top 5%
1.0%
27
Journal of Experimental Botany
219 papers in training set
Top 3%
1.0%
28
Nature Plants
94 papers in training set
Top 2%
1.0%
29
Nature Communications
5641 papers in training set
Top 58%
0.8%
30
Molecular Ecology Resources
171 papers in training set
Top 2%
0.8%