fastCDS: proteome-scale mapping of protein domains to genomic coordinates
Munoz-Esquivel, G.; Fuxman Bass, J. I.; Soto-Ugaldi, L. F.
Show abstract
SummaryMapping protein regions to genomic coordinates underpins the study of exon architecture and the interpretation of clinical variants in their exon context. Existing tools resolve individual queries accurately but scale poorly to proteome-wide analyses. We present fastCDS, a C++ toolkit with command line and Python interfaces for rapid protein-to-genome coordinate mapping from GTF annotations. It matches the accuracy of existing methods while running at least two to three orders of magnitude faster. Mapping all human Pfam domains in seconds, we used the resulting atlas to examine how exonic architecture varies with domain function. Availability and ImplementationfastCDS is freely available under the MIT license at {{https://github.com/SotoLF/fastCDS}} and can be installed with pip install fastCDS or mamba install -c bioconda fastCDS. Pre-built GTF genome indices are archived at Zenodo, DOI: https://zenodo.org/records/21436146. Contactlsoto@rockefeller.edu Supplementary InformationSupplementary data are available at Bioinformatics online.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- AdaLiftOver: High-resolution identification of orthologous regulatory elements with adaptive liftOver 93%
- PROTRIDER: Protein abundance outlier detection from mass spectrometry-based proteomics data with a conditional autoencoder 93%
- hipFG: High-throughput harmonization and integration pipeline for functional genomics data 93%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Helixer--de novo Prediction of Primary Eukaryotic Gene Models Combining Deep Learning and a Hidden Markov Model 93%
- ColabFold - Making protein folding accessible to all 93%
- Genomics 2 Proteins portal: A resource and discovery tool for linking genetic screening outputs to protein sequences and structures 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.