Back

KaMRaT: a C++ toolkit for k-mer count matrix dimension reduction

Xue, H.; Gallopin, M.; Marchet, C.; Nguyen, T. N. H.; Wang, Y.; Bessiere, C.; Gautheret, D.

2024-01-16 bioinformatics
10.1101/2024.01.15.575511 bioRxiv
Show abstract

SummaryKaMRaT is a program for processing large k-mer count tables extracted from high throughput sequencing data. Major functions include scoring k-mers based on count statistics, merging overlapping k-mers into longer contigs and selecting k-mers based on their presence in certain samples. KaMRaT s main application is the reference-free analysis of multi-sample and multi-condition datasets from RNA-seq, as well as ChiP-seq or ribo-seq experiments. KaMRaT enables the identification of condition-specific or differential sequences, irrespective of any gene or transcript annotation. Implementation and availabilityKaMRaT is implemented in C++. Source code and documentation are available via https://github.com/Transipedia/KaMRaT. Container images are available via https://hub.docker.com/r/xuehl/kamrat.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.