From Wikipedia to AI: Measuring 25 years of synthesis of human genetics research in the public-facing information ecosystem
Diaz-Papkovich, A.; Kuntzleman, A.; Davis, S. C.; Ramachandran, S.
Show abstract
Genetics research frequently intersects with ethnicity, nationality, and race, making it uniquely vulnerable to misrepresentation. Yet, 25 years after the initial sequencing of the human genome, there is little understanding of how human genetics research exists in the public-facing information ecosystem. We analyze 3,050,422 historical revisions from 6,738 Wikipedia pages about ethnicity, nationality, and race spanning 25 years. We find genetics terminology is present in 14.8% of these pages (55.5% in the top 1,000 pages) and in 67.8% of pages about nationalities, suggesting research is synthesized to present a biological element to ethnicity and nationality. We also find that 10.1% of 56,908 discussions from these pages contain genetics terminology. We further analyze responses from three popular chatbots queried about nationalities and find that they commonly reference both genetics and Wikipedia. Lastly, we analyze 133 pages from Grokipedia, an AI-generated encyclopedia, and find it mentions genetics more frequently than Wikipedia and hallucinates or misrepresents human genetics research.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Large-scale whole-genome sequencing of three diverse Asian populations in Singapore 92%
- Private information leakage from functional genomics data: Quantification with calibration experiments and reduction via data sanitization protocols 89%
- SPLASH: a statistical, reference-free genomic algorithm unifies biological discovery 89%
Similar papers in this journal
- A simple approach for multiple observations improves power to detect genetic effects and genomic prediction accuracy. 90%
- Inverted genomic regions between reference genome builds in humans impact imputation accuracy and decrease the power of association testing 89%
- Powerful eQTL mapping through low coverage RNA sequencing 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.