P-GRe : an efficient pipeline to maximised pseudogene prediction in plants/eucaryotes
Cabanac, S.; Mathe, C.; Dunand, C.
Show abstract
Formerly considered as part of "junk DNA", pseudogenes are nowadays known for their role in the post-transcriptional regulation of functional genes. In addition, their identification allows a better understanding of gene evolution in the frame of multigenic families. Despite this, there is, to our knowledge, no fully automatic user-friendly software allowing the annotation of pseudogenes on a whole genome. Here, we present Pseudo-Gene Retriever (P-GRe), a fully automated pseudogene prediction software requiring only a genome sequence and its corresponding GFF annotation file. P-GRe detects the sequences of the pseudogenes on a whole genome and returns to the user all their genomic sequences and their pseudo-coding sequences. The ability of P-GRe to finely reconstruct the structure of pseudogenes also allow to obtain a set of proteins virtually encoded by the predicted pseudogenes. We show here that in 70% of the cases, virtual proteins constructed by P-GRe from Arabidopsis thaliana proteome and genome aligned better to their parent protein than their annotated counterpart.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- GeneMark-EP and -EP+: eukaryotic gene prediction with self-training in the space of genes and proteins 96%
- Tailored machine learning models for functional RNA detection in genome-wide screens 94%
- BRAKER2: Automatic Eukaryotic Genome Annotation with GeneMark-EP+ and AUGUSTUS Supported by a Protein Database 94%
Similar papers in this journal
- DARTS: an Algorithm for Domain-Associated RetroTransposon Search in Genome Assemblies 93%
- LSTrAP-Cloud: A User-friendly Cloud ComputingPipeline to Infer Co-functional and RegulatoryNetworks 93%
- The effects of sequence length and composition of random sequence peptides on the growth of E. coli cells 93%
Similar papers in this journal
- High-fidelity (repeat) consensus sequences from short reads using combined read clustering and assembly 94%
- MasterPATH: network analysis of functional genomics screening data 94%
- Comprehensive genome-wide identification of angiosperm upstream ORFs with peptide sequences conserved in various taxonomic ranges using a novel pipeline, ESUCA 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.