Back

Pattern-centric transformation of omics-data sources grounded on multi-wise gene associations aids predictive tasks in TCGA while ensuring interpretability.

Patricio, A.; Costa, R. S.; Henriques, R.

2023-05-30 genomics
10.1101/2023.05.28.542574 bioRxiv
Show abstract

MotivationThe increasing prevalence of omics data sources is pushing the study of regulatory mechanisms underlying complex diseases such as cancer. However, the vast quantities of features produced and the inherent interplay between them lead to a level of complexity that hampers both descriptive and predictive tasks, requiring custom-built algorithms that can extract relevant information from these sources of data. ResultsWe propose a transformation that moves data centered on molecules (e.g. transcripts and proteins) to a new data space focused on putative regulatory modules given by statistically relevant patterns of coexpression. The proposed transformation extracts patterns from the data through biclustering and uses them to create new variables with guarantees of interpretability and discriminative power. The transformation is shown to achieve dimensionality reductions of up to 99% and to increase the predictive performance of various classifiers across multiple omics layers. Our results suggest that a transformation of omics data from gene-centric to pattern-centric data provides benefits to both prediction tasks and human interpretation. The proposed approach is expected to greatly support further bioinformatic analyses for precision medicine applications. AvailabilitySoftware code and the raw results generated are available at github.com/Andrempp/Pattern-Centric-Transformation. Contactandremppatricio@tecnico.ulisboa.pt Supplementary informationSupplementary data are available at Journal Name online.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Artificial Intelligence in Medicine
17 papers in training set
Top 0.1%
12.9%
2
Bioinformatics
1204 papers in training set
Top 2%
11.9%
3
BMC Bioinformatics
457 papers in training set
Top 0.7%
9.7%
4
Frontiers in Genetics
230 papers in training set
Top 0.2%
7.9%
5
PLOS ONE
5266 papers in training set
Top 26%
6.2%
6
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.6%
5.2%
50% of probability mass above
7
Scientific Reports
3612 papers in training set
Top 23%
4.4%
8
Bioinformatics Advances
203 papers in training set
Top 1%
4.0%
9
BMC Genomics
406 papers in training set
Top 2%
3.2%
10
NAR Genomics and Bioinformatics
242 papers in training set
Top 1%
3.2%
11
GigaScience
212 papers in training set
Top 2%
2.4%
12
Genomics
64 papers in training set
Top 0.5%
2.4%
13
Briefings in Bioinformatics
354 papers in training set
Top 4%
2.4%
14
International Journal of Molecular Sciences
494 papers in training set
Top 6%
2.1%
15
PLOS Computational Biology
1863 papers in training set
Top 15%
1.7%
16
Nucleic Acids Research
1281 papers in training set
Top 10%
1.3%
17
PeerJ
308 papers in training set
Top 8%
1.1%
18
Genes
144 papers in training set
Top 3%
1.1%
19
BioData Mining
22 papers in training set
Top 0.8%
0.8%
20
Frontiers in Molecular Biosciences
102 papers in training set
Top 2%
0.8%
21
Heliyon
152 papers in training set
Top 8%
0.8%
22
Journal of Computational Biology
48 papers in training set
Top 1%
0.8%
23
Computational Biology and Chemistry
28 papers in training set
Top 1%
0.6%