Back

Peripheral Blood Nascent RNA Sequencing Captures Ancestry-Linked Differential Gene Expression

Kimm, A.; Uh, E. G.; Kwak, H.

2026-01-15 genomics
10.64898/2026.01.15.699738 bioRxiv
Show abstract

Genetic ancestry reflects inherited population structure that can influence gene regulation and disease susceptibility, yet biomedical studies continue to rely on self-identified race as a proxy for genetic variation. We developed a streamlined framework to infer local and global genetic ancestry directly from nascent RNA sequencing data and tested whether ancestry-based stratification improves detection of transcriptional differences compared with race. Peripheral blood samples from 50 donors were profiled using Peripheral Blood Chromatin Run-On sequencing (pChRO). Ancestry-informative single nucleotide polymorphisms were identified from aligned reads and compared with 1000 Genomes reference populations representing populations in the U.S. using a Bayesian model to infer ancestry locally. Global ancestry proportions were assigned by unsupervised clustering methods. Differential gene expression was analyzed by controlling for age and sex, and the number of true differentially expressed genes (DEGs) was estimated using a q-value framework. Ancestry inference from sparse nascent transcription data recovered heterogeneous local ancestry patterns and separated individuals into two major ancestry clusters, revealing multiple discordances with self-identified race. Grouping by genetically inferred ancestry increased power to detect transcriptional differences relative to to self-identified race, including genes involved in dermatologic and neurodegenerative pathways. These results demonstrate that nascent RNA sequencing enables ancestry inference and that ancestry-based stratification captures transcriptional variation not resolved by race, supporting its integration into population scale functional genomics study.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.