NeighborFinder: an R package inferring local microbial network around a species of interest
Sola, M.; Paravel, A.; Auger, S.; Chatel, J.-M.; Plaza Onate, F.; Le Chatelier, E.; Leclerc, M.; Veiga, P.; Frioux, C.; Mariadassou, M.; Berland, M.
Show abstract
1MotivationUnderstanding interactions from microbiome data is a central aspect in microbial ecology, as it provides insights into ecosystem stability, disease mechanisms, and can be used to design synthetic communities. Current network inference tools reconstruct global networks from co-abundance data, which means they capture the overall correlation structure for the entire set of taxa considered. These approaches are computationally intensive and suboptimal when the focus is on the local neighborhood of specific taxa of interest. ResultsWe introduce NeighborFinder, a local network inference method that enables the targeted discovery of direct neighbors around a species of interest. Using cross-validated multiple linear regression with{ell} 1 penalty and microbiome-specific filters, our approach infers interpretable species-centered interactions, with F1 score [≥] 0.95 on simulated cohorts ranging from 250 to 1000 samples. This method is well-suited for large metagenomic datasets and is particularly valuable for exploratory studies where the targeted hypotheses outweigh the need for global community structure. The approach complements existing methods by being a biologically intuitive and computationally efficient. Availability and ImplementationThe R package is freely available on GitHub: https://github.com/metagenopolis/NeighborFinder. The data and source code used to calculate performances and produce the use case example in this paper can be found respectively at: https://doi.org/10.57745/UPITJ0 and https://doi.org/10.57745/HJLWW4. Supplementary informationSupplementary data are available
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.