Basal Contamination of Sequencing: Lessons from the GTEx dataset
Nieuwenhuis, T. O.; Yang, S.; Verma, R. X.; Pillalamarri, V.; Arking, D.; Rosenberg, A. Z.; McCall, M. N.; Halushka, M. K.
Show abstract
One of the challenges of next generation sequencing (NGS) is read contamination. We used the Genotype-Tissue Expression (GTEx) project, a large, diverse, and robustly generated dataset, to understand the factors that contribute to contamination. We obtained GTEx datasets and technical metadata and validating RNA-Seq from other studies. Of 48 analyzed tissues in GTEx, 26 had variant co-expression clusters of four known highly expressed and pancreas-enriched genes (PRSS1, PNLIP, CLPS, and/or CELA3A). Fourteen additional highly expressed genes from other tissues also indicated contamination. Sample contamination by non-native genes was associated with a sample being sequenced on the same day as a tissue that natively expressed those genes. This was highly significant for pancreas and esophagus genes (linear model, p=9.5e-237 and p=5e-260 respectively). Nine SNPs in four genes shown to contaminate non-native tissues demonstrated allelic differences between DNA-based genotypes and contaminated sample RNA-based genotypes, validating the contamination. Low-level contamination affected 4,497 (39.6%) samples (defined as 10 PRSS1 TPM). It also led [≥] to eQTL assignments in inappropriate tissues among these 18 genes. We note this type of contamination occurs widely, impacting bulk and single cell data set analysis. In conclusion, highly expressed, tissue-enriched genes basally contaminate GTEx and other datasets impacting analyses. Awareness of this process is necessary to avoid assigning inaccurate importance to low-level gene expression in inappropriate tissues and cells.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Common tissue-specific expressions and regulatory mechanisms of c-KIT isoforms with and without GNNK and GNSK sequences across five mammals 93%
- Exogene: A performant workflow for detecting viral integrations from paired-end next-generation sequencing data 92%
- Bioinformatic characterization of angiotensin-converting enzyme 2, the entry receptor for SARS-CoV-2 92%
Similar papers in this journal
- Detecting haplotype-specific transcript variation in long reads with FLAIR2 93%
- CHESS 3: an improved, comprehensive catalog of human genes and transcripts based on large-scale expression data, phylogenetic analysis, and protein structure 93%
- JAFFAL: Detecting fusion genes with long read transcriptome sequencing 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.