Back

Base modification analysis in long read sequencing data using Minimod

Samarasinghe, S.; Deveson, I. W.; Gamaarachchi, H.

2025-07-21 bioinformatics
10.1101/2025.07.16.665072 bioRxiv
Show abstract

Recent advances in long read sequencing technologies have enabled the detection of various DNA and RNA base modifications in addition to standard nucleotide sequences. Both major vendors in this space--Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio)--now include base modification information in their sequencing outputs using MM/ML tags embedded in unaligned BAM files. Each vendor also provides dedicated tools for extracting and analysing these tags, such as ONTs modkit and PacBios pb-CpG-tools. This work presents minimod, a new vendor-agnostic tool designed to extract and analyse any type of base modification from sequencing data generated by any platform that supports MM/ML tags. Benchmarking demonstrates that for DNA data, minimod is ~1.25x faster on a server and ~4x on a laptop compared to modkit and pb-CpG-tools. For RNA data, minimod achieves even greater speedups compared to modkit, ~12x on the server and ~55x on the laptop. Minimod is a free, open-source application written in C and is available at https://github.com/warp9seq/minimod.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.