Back

TransGeneSelector: A Transformer-based Approach Tailored for Key Gene Mining with Small Plant Transcriptomic Datasets

Huang, K.; Tian, J.; Sun, L.; Xie, P.; Zhou, S.; Deng, A.; Mo, P.; Zhou, Z.; Jiang, M.; Li, G.; Wang, Y.; Jiang, X.

2023-09-28 bioinformatics
10.1101/2023.09.26.559592 bioRxiv
Show abstract

Gene mining, particularly from small sample sizes such as in plants, remains a challenge in life sciences. Traditional methods often omit significant genes, while deep learning techniques are hindered by small sample constraints and lack specialized gene mining approaches. This paper presents TransGeneSelector, the first deep learning method tailored for key gene mining in small transcriptomic datasets, ingeniously integrating data augmentation, sample filtering, and a Transformer-based classifier. Tested on Arabidopsis thaliana seeds germination classification using just 79 samples, it not only achieves classification performance on par with, if not superior to, Random Forest and SVM but also excels in identifying upstream regulatory genes that Random Forest might miss, and these pinpointed genes more accurately reflect the metabolic processes inherent in seed germination. TransGeneSelectors ability to mine vital genes from limited datasets signifies its potential as the current state-of-the-art in gene mining in small sample scenarios, providing an efficient and versatile solution for this critical research area.

Matching journals

The top 11 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.