Back

Revaluation of old data with new techniques reveals novel insights into the celiac microbiome

Colgan, J. J.; Burns, M. B.

2022-10-07 bioinformatics
10.1101/2022.10.05.510990 bioRxiv
Show abstract

Celiac disease is an autoimmune disorder of the small intestine in which gluten, an energy-storage protein expressed by wheat and other cereals, elicits an immune response leading to villous atrophy. Despite a strong genetic component, the disease arises sporadically throughout life, leading us to hypothesize the the microbiome might be a trigger for celiac disease. Here, we took microbiome data from 3 prior studies examining celiac disease and the microbiome and analyzed this data with newer computational tools and databases: the dada2 and PICRUSt2 pipelines and the SILVA database. Our results both confirmed findings of previous studies and generated new data regarding the celiac microbiome of India and Mexico. Our results showed that, while some aspects of prior reports are robust, older datasets must be reanalyzed with new tools to ascertain which findings remain accurate while also uncovering new findings. IMPORTANCEBioinformatics is a rapidly developing field, with new computational tools released yearly. It is thus important to revisit results generated using older tools to determine whether they are also revealed by currently available technology. Celiac disease is an autoimmune disorder that affects up to 2% of the worlds population. While the ultimate cause of celiac disease is unknown, many researchers hypothesize that changes to the intestinal microbiome play a role in the diseases progression. Here, we have re-analyzed 16S rRNA data from several previous celiac studies to determine whether previous results are also uncovered using new computational tools.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.