Assessing NGS-based computational methods for predicting transcriptional regulators with query gene sets
Lu, Z.; Xiao, X.; Zheng, Q.; Wang, X.; Xu, L.
Show abstract
This article provides an in-depth review of computational methods for predicting transcriptional regulators with query gene sets. Identification of transcriptional regulators is of utmost importance in many biological applications, including but not limited to elucidating biological development mechanisms, identifying key disease genes, and predicting therapeutic targets. Various computational methods based on next-generation sequencing (NGS) data have been developed in the past decade, yet no systematic evaluation of NGS-based methods has been offered. We classified these methods into two categories based on shared characteristics, namely library-based and region-based methods. We further conducted benchmark studies to evaluate the accuracy, sensitivity, coverage, and usability of NGS-based methods with molecular experimental datasets. Results show that BART, ChIP-Atlas, and Lisa have relatively better performance. Besides, we point out the limitations of NGS-based methods and explore potential directions for further improvement. Key pointsO_LIAn introduction to available computational methods for predicting functional TRs from a query gene set. C_LIO_LIA detailed walk-through along with practical concerns and limitations. C_LIO_LIA systematic benchmark of NGS-based methods in terms of accuracy, sensitivity, coverage, and usability, using 570 TR perturbation-derived gene sets. C_LIO_LINGS-based methods outperform motif-based methods. Among NGS methods, those utilizing larger databases and adopting region-centric approaches demonstrate favorable performance. BART, ChIP-Atlas, and Lisa are recommended as these methods have overall better performance in evaluated scenarios. C_LI
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Identification of upstream transcription factor binding sites in orthologous genes using mixed Student's t-test statistics 96%
- Reconstruction Set Test (RESET): a computationally efficient method for single sample gene set testing based on randomized reduced rank reconstruction error 95%
- Mcadet: a feature selection method for fine-resolution single-cell RNA-seq data based on multiple correspondence analysis and community detection 95%
Similar papers in this journal
- Coffee: Consensus Single Cell-Type Specific Inference For Gene Regulatory Networks 95%
- PTFSpot: Deep co-learning on transcription factors and their binding regions attains impeccable universality in plants 94%
- A novel splicing graph allows a direct comparison between exon-based and splice junction-based approaches to alternative splicing detection 94%
Similar papers in this journal
Similar papers in this journal
- Evidence for the role of transcription factors in the co-transcriptional regulation of intron retention 95%
- Robustness and applicability of functional genomics tools on scRNA-seq data 95%
- Simultaneous smoothing and detection of topological units of genome organization from sparse chromatin contact count matrices with matrix factorization 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.