Back

Synthesizing human-specific spike-in standards for RNA-seq experiments and assessing their technical performance

Qin, R.; Fan, W.; Ding, F.; Wang, S.; He, B.; Hou, M.; Lin, Q.; Cui, P.; Liu, W.

2025-04-07 bioinformatics
10.1101/2025.04.01.646725 bioRxiv
Show abstract

Using spike-in standards for RNA-seq experiments is critical to evaluate technical bias introduced during sample preparation, sequencing and analysis. Although some external RNA spike-in standards have been developed, species-specific spike-in standard was not reported yet. Here we developed the human-specific spike-in standards with 65 controls. We first extracted human endogenous RNAs with various lengths and GC contents, introduced random mutations approximately every 75 bp in each RNA, and then synthesized these RNAs as spike-in RNAs. After that, four mixtures of these spike-in RNAs covering a 220 dynamic range were obtained. To ensure the accuracy of RNA concentration, two rounds of ddPCR were conducted for each spike-in RNA and the intraclass correlation coefficient between two ddPCRs ranged from 0.9954 to 0.9971 after removing the two spike-in RNAs with the largest concentration difference. Furthermore, we showed that the sequencing error profiles were distinct between platforms and the library preparation procedures were related with the discrepancies in spike-in RNA read percent, transcript abundance, sequence coverage distribution, and differential gene expression. In addition, two regression models of sequence coverage were built based on RNA second structure and GC content, and 86.62%-91.78% of the variation can be explained. Our study demonstrates the technical performance of the human-specific spike-in standards for RNA-seq experiments and illustrates the biases across different libraries, platforms, and laboratories.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above