Back

Ensemble Learning: Predicting Human Pathogenicity of Hematophagous Arthropod Vector-Borne Viruses

Hu, H.; Zhao, C.; Jin, M.; Chen, J.; Liu, X.; Guo, J.; Shi, H.; Wang, C.; Chen, Y.

2023-12-31 public and global health
10.1101/2023.12.30.23300660 medRxiv
Show abstract

Hematophagous arthropods serve as crucial vectors for numerous viruses, posing significant public health risks due to their potential for zoonotic spillover. Despite the advances in metagenomics expanding our understanding of arbovirus diversity, traditional phylogenetic approaches often miss the pathogenic potential of viruses not yet identified in humans. Here, we curated two datasets: one with 294 viruses and 36 epidemiological characteristics (including virus properties, vector hosts, and non-vector hosts), and another with 71,622 viral sequences focusing on pathogenic traits. Using these datasets, we developed a regression model and a prediction model to assess and predict viral pathogenicity. Using these datasets, we developed a regression model and a prediction model to assess and predict viral pathogenicity. Our regression model (R2 = 90.6%) reveals a strong correlation between non-vector host diversity, especially within Perissodactyla and Carnivora orders, and virus pathogenicity. The prediction model (F1 score = 96.79%) identifies key pathogenic functions such as "Viral adhesion" and "Host xenophagy" as enhancers of pathogenic potential, while the "Viral invasion" function was associated with an inverse effect. Validation against an external independent dataset confirmed the models ability to identify pathogenic viruses and revealed the potential threat posed by Palma and Zaliv Terpeniya viruses, previously undetected in humans. These findings highlight the necessity of integrating predictive models with metagenomic data to provide early warnings of potential zoonotic viruses carried by hematophagous vectors at the strain level, enhancing public health responses and preparedness.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.