Back

pyfraglib: An integrated cfDNA fragmentomics platform

Schuette, D.; Godfrey, L. K.; Schneider, J.; Borchmann, S.; Heger, J.-M.; Schwarz, R. F.

2026-07-24 bioinformatics
10.64898/2026.07.22.740152 bioRxiv
Show abstract

SummaryCell-free DNA (cfDNA) fragmentomics is the analysis of a diverse set of cfDNA fragment features, e.g. fragment length profiles, windowed protection scores, and end motifs. As such it requires software tooling for fragment extraction, statistical feature modeling, and cohort-level comparative analysis. In silico simulations can facilitate the development and validation of new methods by generating testing datasets with known ground truth. Existing tools address individual aspects of this workflow but none provide all necessary capabilities within a single package. ResultsWe present pyfraglib, a platform integrating fragment extraction from short- and long-read sequencing, statistical feature modeling (Gaussian mixture and NMF decomposition of fragment length profiles, end motif diversity, windowed protection scores), cohort-level differential testing of said features, and a simulation module. The library is exposed through a command-line interface, a Python API, and a Nextflow pipeline. We demonstrate pyfraglib in two ways. First, on two simulated 20-sample cohorts we show that pyfraglibs per-sample and cohort-level analyses recover the differences introduced by construction. Second, we apply pyfraglib to 89 cfDNA samples from a central nervous system lymphoma (CNSL) study and construct a fragmentomics score combining an NMF signature with end motif and WPS summaries via a classifier trained on cerebrospinal fluid and healthy donor plasma samples. As a proof of concept and applied to patient plasma samples, the score identifies a high-risk subgroup with worse failure-free survival (log-rank p=0.0205). Conclusionspyfraglib integrates sample- and cohort-level fragmentomics analyses as well as in silico simulation within a consistently engineered Python framework. pyfraglib source code and documentation are available at https://github.com/ICCB-Cologne/pyfraglib.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.