Back

A framework for Polinton-like virus diversity across aquatic microbiomes reveals links to multiple viral classes and Nucleocytoviricota

Bellas, C.; Sommaruga, R.

2026-06-19 microbiology
10.64898/2026.06.19.733378 bioRxiv
Show abstract

Polinton-like viruses (PLVs) are among the most abundant eukaryotic DNA viruses in aquatic environments. Despite their extensive diversity, broad host range and variable gene content, they are commonly treated as a single group, which obscures their evolutionary relationships and complicates their classification. Through analysing thousands of viral genomes from aquatic ecosystems and public metagenomic datasets, we clarify the evolutionary structure encompassed by the term PLV. Using sensitive profile Hidden Markov Model (HMM) comparisons, phylogenies of conserved capsid morphogenetic genes and gene content analysis, we show that viruses referred to as PLVs are distributed across multiple deep lineages spanning at least three currently recognised viral classes. These include the Gosseviruses, aquatic viruses related to Maverick-Polintons in animal genomes. They also include a continuum of related viruses from 15 kb PLVs to the 45 kb Mriyaviruses and more broadly, to the Nucleocytoviricota, potentially representing extant relatives of giant viruses. Our findings suggest that PLVs do not fit neatly within existing taxonomic boundaries, reflecting a complex history of horizontal gene transfer and diversification of life strategies. To support future discovery, we provide a curated set of HMMs representing the known capsid diversity of PLVs, Maverick-Polintons, and virophages. This toolkit enables sensitive detection and identification of PLVs across metagenomic and eukaryotic genome datasets. Our study provides an evolutionary framework for interpreting PLV diversity and a foundation for future refinement of their classification.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.