Back

Statistics used by ecologists: the rise of R statistical software, GLM(M)s and Networks

Silva, E. A.

2025-12-09 ecology
10.64898/2025.12.04.692382 bioRxiv
Show abstract

About 15 years ago, statistical softwares were rarely employed to conduct statistical analyses, and biologists/ecologists used roughly 100 statistical procedures in research. This number has likely increased substantially with the development of new analytical methods and the widespread adoption of computers and softwares. In this study, I investigated the temporal variation in statistical procedures and the use of softwares used in studies on plant-ant interactions. Data were collected from 142 published papers covering a period between 1979 to 2023. Information related to statistics terminology, the softwares cited and R software packages were retrieved. Each paper had on average eight statistical procedures and there was a significant increase in procedures over the years. The R software was cited in almost 80% of studies, surpassing by far the other softwares. Classical analyses such as t-tests, correlations and chi-squares are still used in research; nonetheless, from 2012 onward a new set of analyses including GLM, GLMM and Networks have emerged and are being used as the leading analyses in studies of plant-ant interactions. This coincides with the adoption of R software by scientists in this field. The most used R packages were vegan, lme4 and bipartite, followed by several other related to linear (mixed) models. Statistics literacy is a mandatory aspect of research in plant-ant interactions. The increase in statistical procedures over the years, as well as the emergence of new tests can be considered as a natural evolution of the field with demand for robustness in methods and analyses.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.