TD2: finding protein coding regions in transcripts
Mao, A.; Ji, H. J.; Haas, B.; Salzberg, S.; Sommer, M. J.
Show abstract
The transcriptome encompasses all RNA transcripts in eukaryotic cells, orchestrating gene expression and regulating cellular function, development, and adaptation. Identifying open reading frames (ORFs) in transcripts is a critical step in transcriptome analysis. We introduce TD2, a new tool for ab initio annotation of protein-coding ORFs in transcripts. We find TD2 to be sensitive and precise when compared to other state-of-the-art tools in reference transcripts and transcriptome assemblies from a diverse array of eukaryotes. TD2 is available at https://github.com/Markusjsommer/TD2. The project is open-source, developed in Python with PyTorch, and is freely available to all academic, government, and commercial users under the MIT license.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Tailored machine learning models for functional RNA detection in genome-wide screens 96%
- GeneMark-EP and -EP+: eukaryotic gene prediction with self-training in the space of genes and proteins 95%
- FLYNC: A Machine Learning-Driven Framework for Discovering Long Non-Coding RNAs in Drosophila melanogaster 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.