Back

Improvement in Neoantigen Prediction via Integration of RNA Sequencing Data for Variant Calling

Nguyen, B. Q. T.; Tran, T. P. D.; Nguyen, H. T.; Nguyen, T. N.; Pham, T. M. Q.; Nguyen, H. T. P.; Tran, D. H.; Nguyen, T. T. V.; Tran, T. S.; Pham, T.-V. N.; Le, M.-T.; Phan, M.-D.; Giang, H.; Nguyen, H.-N.; Tran, L. S.

2023-07-03 immunology
10.1101/2023.07.02.547404 bioRxiv
Show abstract

Neoantigen-based immunotherapy has emerged as a promising strategy for improving the life expectancy of cancer patients. This therapeutic approach heavily relies on accurate identification of cancer mutations using DNA sequencing (DNAseq) data. However, current workflows tend to provide a large number of neoantigen candidates, of which only a limited number elicit efficient and immunogenic T-cell responses suitable for downstream clinical evaluation. To overcome this limitation and increase the number of high-quality immunogenic neoantigens, we propose integrating RNA sequencing (RNAseq) data into the mutation identification step in the neoantigen prediction workflow. In this study, we characterize the mutation profiles identified from DNAseq and/or RNAseq data in tumor tissues of 25 patients with colorectal cancer (CRC). We detected only 22.4% of variants shared between the two methods. In contrast, RNAseq-derived variants displayed unique features of affinity and immunogenicity. We further established that neoantigen candidates identified by RNAseq data significantly increased the number of highly immunogenic neoantigens (confirmed by ELISpot) that would otherwise be overlooked if relying solely on DNAseq data. In conclusion, this integrative approach holds great potential for improving the selection of neoantigens for personalized cancer immunotherapy, ultimately leading to enhanced treatment outcomes and improved survival rates for cancer patients.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.