Statistical knockoffs improve biomarker discovery fromtranscriptomic data
CARTIER, J.; LAGOAS, J.; FERMANIAN, A.; Azencott, C.-A.; MASSIP, F.
Show abstract
Advances in sequencing technologies have enabled the generation of large amounts of data, offering new possibilities to identify relationships between biological units (e.g. genes) and phenotypic traits (e.g. disease outcomes). Yet, identifying these associations using variable selection methods remains challenging due to the high dimension (p >> n) and the correlation structure of the data. To address these challenges, we study the applicability of the knockoff (KO) procedure. Introduced by Barber and Candes in 2015, the KO variable selection procedure has shown promising results on real biological data, such as Genome Wide Association Studies. This method seeks to identify the truly important predictors by overcoming the correlation structure between variables while controlling the false discovery rate. Here, we study the applicability of the KO procedure on transcriptomic data in a classification setting. We conduct an extensive simulation study using real transcriptomic data to evaluate the performance of the KO framework in the context of high-dimensional classification. We find that the KO framework outperforms widely used variable selection models, and that using KO aggregation to mitigate the effect of KO stochasticity improves stability while maintaining the same power. Finally applied to three real transcriptomic datasets, the KO framework made very few discoveries, highlighting its conservative nature and suggesting that other methods may substantially overestimate the number of relevant features.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Biological networks and GWAS: comparing and combining network methods to understand the genetics of familial breast cancer susceptibility in the GENESIS study 96%
- RCFGL: Rapid Condition adaptive Fused Graphical Lasso and application to modeling brain region co-expression networks 96%
- Application of Modular Response Analysis to Medium- to Large-Size Biological Systems 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.