LncPlankton V1.0: a comprehensive collection of plankton long non-coding RNAs
Debit, A.; Vincens, P.; Bowler, C.; Cruz de Carvalho, H.
Show abstract
Long considered as transcriptional noise, long non-coding RNAs (lncRNAs) are emerging as central, regulatory molecules in a multitude of eukaryotic species, from plants to animals to fungi. Yet, our knowledge about the occurrence of these molecules in the marine environment, namely in planktonic protists, is still elusive. To fill this gap of knowledge we developed LncPlankton v1.0, which is the first comprehensive database of marine plankton lncRNAs. By integrating the predictions derived from ten distinctive coding potential prediction tools in a majority voting setting, we identified 2,210,359 lncRNAs distributed across 414 marine plankton species from over nine different phyla. A user-friendly, open-access web interface for the exploration of the database was implemented (https://www.lncplankton.bio.ens.psl.eu/). We believe LncPlankton v1.0 will serve as a rich resource for studies of lncRNAs that will contribute to small- and large-scale analyses in a wide range of marine plankton species and allow comparative analysis well beyond the marine environment.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Kmerator Suite: design of specific k-mer signatures andautomatic metadata discovery in large RNA-Seq datasets. 96%
- FLYNC: A Machine Learning-Driven Framework for Discovering Long Non-Coding RNAs in Drosophila melanogaster 96%
- miRge3.0: a comprehensive microRNA and tRF sequencing analysis pipeline 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- RNAlysis: analyze your RNA sequencing data without writing a single line of code 92%
- Taxonomy-aware, sequence similarity ranking reliably predicts phage-host relationships 91%
- Single-cell Mayo Map (scMayoMap): an easy-to-use tool for cell type annotation in single-cell RNA-sequencing data analysis 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.