Back

TD2: finding protein coding regions in transcripts

Mao, A.; Ji, H. J.; Haas, B.; Salzberg, S.; Sommer, M. J.

2025-04-14 genomics
10.1101/2025.04.13.648579 bioRxiv
Show abstract

The transcriptome encompasses all RNA transcripts in eukaryotic cells, orchestrating gene expression and regulating cellular function, development, and adaptation. Identifying open reading frames (ORFs) in transcripts is a critical step in transcriptome analysis. We introduce TD2, a new tool for ab initio annotation of protein-coding ORFs in transcripts. We find TD2 to be sensitive and precise when compared to other state-of-the-art tools in reference transcripts and transcriptome assemblies from a diverse array of eukaryotes. TD2 is available at https://github.com/Markusjsommer/TD2. The project is open-source, developed in Python with PyTorch, and is freely available to all academic, government, and commercial users under the MIT license.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.