Back

An Integrated Deep Learning Framework for Small-Sample Biomedical Data Classification: Explainable Graph Neural Networks with Data Augmentation for RNA sequencing Dataset

Guler, F.; Goksuluk, D.; Xu, M.; Choudhary, G.; agraz, m.

2026-02-24 genetic and genomic medicine
10.64898/2026.02.22.26346827 medRxiv
Show abstract

Applying deep learning models to RNA-Seq data poses substantial challenges, primarily due to the high dimensionality of the data and the limited sample sizes. To address these issues, this study introduces an advanced deep learning pipeline that integrates feature engineering with data augmentation. The engineering application focuses on biomedical engineering, specifically the classification of RNA-Seq datasets for disease diagnosis. The proposed framework was initially validated on synthetic datasets generated from Naive Bayes, where MLP-based augmentation yielded a notable improvement in predictive performance. Building on this foundation, we applied the approach to chromophobe renal cell carcinoma (KICH) RNA-Seq data from The Cancer Genome Atlas (TCGA). Following standard preprocessing steps normalization, transformation, and dimensionality reduction, the analysis concentrated on three main aspects: augmentation strategies, preprocessing methods, and explainable AI (XAI) techniques in relation to classification outcomes. Feature selection was performed through PCA, Boruta, and RF-based methods. Three augmentation strategies linear interpolation, SMOTE, and MixUp were evaluated. To maintain methodological rigor, augmentation was applied exclusively to the training set, while the test set was held out for unbiased evaluation. Within this framework, we conducted a comparative assessment of multiple deep learning architectures, including MLP, GNN, and the recently proposed Kolmogorov-Arnold networks (KAN). The GNN achieved the highest classification accuracy (99.47%) when trained with MixUp augmentation combined with RF feature selection, and achieved the best F1 score (0.9948). Consequently, the GNN-based XAI framework was applied to the RF dataset enriched with MixUp. XAI analyses identified the top 20 most influential genes, such as HNF4A, DACH2, MAPK15, and NAT2, which played the greatest role in classification, thereby confirming the biological plausibility of the model outputs. To further validate model robustness, cervical cancer and Alzheimers RNA-Seq datasets were also tested, yielding consistent and reliable results. Overall, the findings highlight the value of incorporating data augmentation into deep learning models for RNA-Seq analysis, not only to improve predictive performance but also to enhance biological interpretability through explainable AI approaches.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

1
Frontiers in Bioinformatics
49 papers in training set
Top 0.1%
9.6%
2
Scientific Reports
3612 papers in training set
Top 8%
7.7%
3
Bioinformatics
1204 papers in training set
Top 4%
6.1%
4
GigaScience
212 papers in training set
Top 0.6%
5.4%
5
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.6%
5.4%
6
Frontiers in Genetics
230 papers in training set
Top 0.5%
5.4%
7
NAR Genomics and Bioinformatics
242 papers in training set
Top 1%
4.2%
8
Briefings in Bioinformatics
354 papers in training set
Top 2%
4.0%
9
Bioinformatics Advances
203 papers in training set
Top 2%
3.4%
50% of probability mass above
10
Computers in Biology and Medicine
128 papers in training set
Top 1%
3.2%
11
BMC Bioinformatics
457 papers in training set
Top 3%
3.2%
12
IEEE/ACM Transactions on Computational Biology and Bioinformatics
38 papers in training set
Top 0.3%
3.1%
13
Heliyon
152 papers in training set
Top 2%
2.6%
14
PLOS ONE
5266 papers in training set
Top 43%
2.4%
15
Nature Communications
5641 papers in training set
Top 41%
2.3%
16
Informatics in Medicine Unlocked
22 papers in training set
Top 0.5%
1.9%
17
Journal of Translational Medicine
57 papers in training set
Top 0.8%
1.7%
18
Communications Medicine
113 papers in training set
Top 2%
1.7%
19
PLOS Computational Biology
1863 papers in training set
Top 14%
1.7%
20
BMC Genomics
406 papers in training set
Top 5%
1.7%
21
International Journal of Molecular Sciences
494 papers in training set
Top 8%
1.7%
22
Patterns
78 papers in training set
Top 1%
1.6%
23
Artificial Intelligence in the Life Sciences
13 papers in training set
Top 0.1%
1.6%
24
Cancers
213 papers in training set
Top 3%
1.3%
25
Journal of the American Medical Informatics Association
71 papers in training set
Top 2%
1.0%
26
iScience
1154 papers in training set
Top 32%
0.9%
27
Journal of Biomedical Informatics
47 papers in training set
Top 1%
0.8%
28
Frontiers in Bioengineering and Biotechnology
98 papers in training set
Top 3%
0.8%
29
Nucleic Acids Research
1281 papers in training set
Top 14%
0.8%
30
Artificial Intelligence in Medicine
17 papers in training set
Top 0.9%
0.6%