Back

The Portable Microhaplotype Object and Tools

Hathaway, N. J.; Murie, K.; Murphy, M.; Simkin, A.; Amaya-Romero, J.; Hubbard, A.; Briggs, J.; Aranda-Diaz, A.; Early, A. M.; Wesolowski, A.; Neafsey, D. E.; Bailey, J. A.; Greenhouse, B.

2025-12-12 bioinformatics
10.64898/2025.12.10.693568 bioRxiv
Show abstract

Structured AbstractO_ST_ABSMotivationC_ST_ABSThe rapid increase in the generation of targeted sequencing data offers immense potential for research, medicine, and public health, however the lack of an established standard for these data has led to disparate solutions for data storage. A widely accepted standard is essential for data sharing, reuse, and the coordinated development of interoperable analysis tools. ResultsWe propose the Portable Microhaplotype Object (PMO), a standardized format for efficiently and losslessly storing phased targeted sequencing data (microhaplotypes). The PMO format is JSON-based, allowing efficient, relational storage of genetic data together with relevant metadata to minimize orphaned data. The format includes required fields and a curated set of optional fields leveraging established ontologies. To facilitate ease of use, we developed pmotools-python, an open-source package for creating, manipulating, and exporting PMO data into common formats. Additionally, we provide a simple web-based app to quickly create PMO files from tabular inputs, making the format accessible to a wide variety of users. Example datasets from Plasmodium, Anopheles, Escherichia coli, and Staphylococcus aureus demonstrate the broad applicability of the approach. PMO will streamline data sharing, foster interoperability, and accelerate the development of harmonized analysis tools. Availability and implementationThe Portable Microhaplotype Object (PMO) project, including the ontology specification, software tools, example datasets, and tutorials, is freely available at https://plasmogenepi.github.io/PMO_Docs/. Key software components and datasets have archived releases with DOIs to ensure permanence, detailed in the Supplementary Text 1-5.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.