Back

Toward Best Practice in Identifying Subtle Differential Expression with RNA-seq: A Real-World Multi-Center Benchmarking Study Using Quartet and MAQC Reference Materials

Wang, D.; Liu, Y.; Zhang, Y.; Chen, Q.; Han, Y.; Hou, W.; Liu, C.; Yu, Y.; Li, Z.; Li, Z.; Zhao, J.; Zheng, Y.; Shi, L.; Li, J.; Zhang, R.

2023-12-10 genomics
10.1101/2023.12.09.570956 bioRxiv
Show abstract

Translating RNA-seq into clinical diagnostics requires ensuring the reliability of detecting clinically relevant subtle differential expressions, such as those between different disease subtypes or stages. Moreover, cross-laboratory reproducibility and consistency under diverse experimental and bioinformatics workflows urgently need to be addressed. As part of the Quartet project, we presented a comprehensive RNA-seq benchmarking study utilizing Quartet and MAQC RNA reference samples spiked with ERCC controls in 45 independent laboratories, each employing their in-house RNA-seq workflows. We assessed the data quality, accuracy and reproducibility of gene expression and differential gene expression and compared over 40 experimental processes and 140 combined differential analysis pipelines based on multiple ground truths. Here we show that real-world RNA-seq exhibited greater inter-laboratory variations when detecting subtle differential expressions between Quartet samples. Experimental factors including mRNA enrichment methods and strandedness, and each bioinformatics step, particularly normalization, emerged as primary sources of variations in gene expression and have a more pronounced impact on the subtle differential expression measurement. We underscored the pivotal role of experimental execution over the choice of experimental protocols, the importance of strategies for filtering low-expression genes, and optimal gene annotation and analysis tools. In summary, this study provided best practice recommendations for the development, optimization, and quality control of RNA-seq for clinical diagnostic purposes.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.