A Golden Standard for Transcription Factor-Gene Regulatory Interactions in Escherichia coli K-12: How reliable are they based on the methods supporting them?
Lara, P.; Gama-Castro, S.; Salgado, H.; Rioualen, C.; Muniz- Rascado, L. J.; Garcia-Sotelo, J. S.; Tierrafria, V. H.; Collado-Vides, J.
Show abstract
Post-genomic implementations have expanded the experimental strategies to identify elements involved in the regulation of transcription initiation. As new methodologies emerge, a natural step is to compare their results with those from established methodologies, such as the classic methods of molecular biology used to characterize transcription factor binding sites, promoters, or transcription units. In the case of Escherichia coli K-12, the best-studied microorganism, for the last 30 years we have continuously gathered such knowledge from original scientific publications, and have organized it in two databases, RegulonDB and EcoCyc. Furthermore, since RegulonDB version 11.0 (1), we offer comprehensive datasets of binding sites from chromatin immunoprecipitation combined with sequencing (ChIP-seq), ChIP combined with exonuclease digestion and next-generation sequencing (ChIP-exo), genomic SELEX screening (gSELEX), and DNA affinity purification sequencing (DAP-seq) HT technologies, as well as additional datasets for transcription start sites, transcription units and RNA sequencing (RNA-seq) expression profiles. Here, we present for the first time an analysis of the sources of knowledge supporting the collection of transcriptional regulatory interactions (RIs) of E. coli K-12. An RI is formed by the transcription factor, its positive or negative effect on a promoter, a gene or transcription unit. We improved the evidence codes so that the specific methods are described, and we classified them into seven independent groups. This is the basis for our updated computation of confidence levels, weak, strong, or confirmed, for the collection of RIs. We compare the confidence levels of the RI collection before and after adding HT evidence illustrating how knowledge will change as more HT data and methods appear in the future. Users can generate subsets filtering out the method they want to benchmark and avoid circularity, or keep for instance only the confirmed interactions. The comparison of different HT methods with the available datasets indicate that ChIP-seq recovers the highest fraction (>70%) of binding sites present in RegulonDB followed by gSELEX, DAP-seq and ChIP-exo. There is no other genomic database that offers this comprehensive high-quality anatomy of evidence supporting a corpus of transcriptional regulatory interactions.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genomic background sequences systematically outperform synthetic ones in de novo motif discovery for ChIP-seq data 95%
- DecoPath: A web application for decoding pathway enrichment analysis 94%
- Poly-Enrich: Count-based Methods for Gene Set Enrichment Testing with Genomic Regions and Updates to ChIP-Enrich 93%
Similar papers in this journal
- Identifying promoter sequence architectures via a chunking-based algorithm using non-negative matrix factorisation 94%
- Identification of upstream transcription factor binding sites in orthologous genes using mixed Student's t-test statistics 93%
- Computational identification and experimental characterization of preferred downstream positions in human core promoters 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.