elDORS: An elevated Database Of RNA Sequences
Dutta, N.; Vicens, Q.
Show abstract
Massive and integrated sequence databases have revolutionized computational protein structure prediction. RNA lags due to a lack of consolidated sequence resources. To bridge this gap, we developed elDORS_raw, which comprises up-to-date sequence information, including recent metagenomes and transcriptome sequence data. We optimized an 80% sequence-identity clustered version named elDORS for use with the RNAcmap3 split-strategy for multiple sequence alignment (MSA) and the widely used rMSA pipeline. Our benchmarking demonstrates that elDORS-augmented MSA pipelines match or exceed the alignment depth obtained with massive legacy databases across different queries, including blind CASP challenges, effectively eliminating sequence retrieval failures challenging orphan RNAs. To aid homology searches, predicting RNA properties, training new models, and other downstream tasks, elDORS is freely accessible. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=197 HEIGHT=200 SRC="FIGDIR/small/737016v1_ufig1.gif" ALT="Figure 1"> View larger version (57K): org.highwire.dtl.DTLVardef@d7549aorg.highwire.dtl.DTLVardef@f36ab1org.highwire.dtl.DTLVardef@e1a07eorg.highwire.dtl.DTLVardef@efdab4_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 5 journals account for 50% of the predicted probability mass.