The GENCODE CLS project: massively expanding the lncRNA catalog through capture long-read RNA sequencing
Perteghella, T.; Kaur, G.; Carbonell-Sala, S.; Gonzalez-Martinez, J.; Hunt, T.; Madry, T.; Jungreis, I.; Degalez, F.; Arnan, C.; Nurtdinov, R.; Lagarde, J.; Borsari, B.; Sisu, C.; Jiang, Y.; Bennett, R.; Berry, A.; Blangiewicz, M.; Cerdan-Velez, D.; Cochran, K.; Vara, C.; Davidson, C.; Donaldson, S.; Dursun, C.; Gonazlez-Lopez, S.; Gopal Das, S.; Lawrence, K.; Nachun, D.; Hardy, M.; Hollis, Z.; Kay, M.; Montanes, J. C.; Ni, P.; Palumbo, E.; Pulido-Quetglas, C.; Suner, M.-M.; Yu, X.; Zhang, D.; Aguet, F.; Ardlie, K.; Montgomery, S. B.; Loveland, J. E.; Alba, M. M.; Diekhans, M.; Tanzer, A.; Mud
Show abstract
Accurate and complete gene annotations are indispensable for understanding how genome sequences encode biological functions. For more than twenty years, the GENCODE consortium has developed reference annotations for the human and mouse genomes, becoming a foundation for biomedical and genomics communities worldwide. Nevertheless, collections of important yet poorly-understood gene classes like long non-coding RNAs (lncRNAs) remain incomplete and scattered across multiple, uncoordinated catalogs. To address this, GENCODE has undertaken the most comprehensive lncRNA annotation effort to date. This is founded on the manually supervised computational annotation of full-length targeted long-read sequencing, on matched embryonic and adult tissues, of orthologous regions in human and mouse. Altogether 17,931 human genes (140,268 transcripts) and 22,784 mouse genes (136,169 transcripts) have been added to the GENCODE catalog representing a 2-fold and 6-fold growth in transcripts, respectively - the greatest increase in the number of annotated human genes since the sequencing of the human genome. Our targeted design assigned human-mouse orthologs at a rate beyond previous studies, tripling the number of human disease-associated lncRNAs that have mouse orthologs. Novel lncRNA genes consistently exhibit biological signals of functionality, and they greatly enhance the functional interpretability of the human genome. While poorly expressed in bulk RNA-Seq samples, many of them are highly expressed in specific cell populations, maybe even contributing to cell-type determination. The expanded GENCODE lncRNA annotations mark a critical step toward deciphering the human and mouse genomes.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Uncovering functional lncRNAs by scRNA-seq with ELATUS 98%
- A spatial long-read approach at near-single-cell resolution reveals developmental regulation of splicing and polyadenylation sites in distinct cortical layers and cell types. 97%
- Identification and analysis of splicing quantitative trait loci across multiple tissues in the human genome 97%
Similar papers in this journal
- Normal and cancer tissues are accurately characterised by intergenic transcription at RNA polymerase 2 binding sites 98%
- Variant-resolved prediction of context-specific isoform variation with a graph-based attention model 97%
- Impact of disease-associated chromatin accessibility QTLs across immune cell types and contexts 96%
Similar papers in this journal
- Transcriptional kinetics and molecular functions of long non-coding RNAs 97%
- Dynamic network-guided CRISPRi screen reveals CTCF loop-constrained nonlinear enhancer-gene regulatory activity in cell state transitions 97%
- Tissue-specific enhancer-gene maps from multimodal single-cell data identify causal disease alleles 96%
Similar papers in this journal
- Systematic assessment of long-read RNA-seq methods for transcript identification and quantification 97%
- A systematic benchmark of Nanopore long read RNA sequencing for transcript level analysis in human cell lines 97%
- Multiplexed spatial mapping of chromatin features, transcriptome, and proteins in tissues 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.