A rigorous benchmarking of alignment-based HLA callers for RNA-seq data
Yu, D.; Ayyala, R.; Mangul, S.
Show abstract
Accurate identification of human leukocyte antigen (HLA) alleles is essential for various clinical and research applications, such as transplant matching and drug sensitivities. Recent advances in RNA-seq technology have made it possible to impute HLA types from sequencing data, spurring the development of a large number of computational HLA typing tools. However, the relative performance of these tools is unknown, limiting the ability for clinical and biomedical research to make informed choices regarding which tools to use. Here we report the study design of a comprehensive benchmarking of the performance of 12 HLA callers across 682 RNA-seq samples from 8 datasets with molecularly defined gold standard at 5 loci, HLA-A, -B, -C, -DRB1, and -DQB1. For each HLA typing tool, we will comprehensively assess their accuracy, compare default with optimized parameters, and examine for discrepancies in accuracy at the allele and loci levels. We will also evaluate the computational expense of each HLA caller measured in terms of CPU time and RAM. We also plan to evaluate the influence of read length over the HLA region on accuracy for each tool. Most notably, we will examine the performance of HLA callers across European and African groups, to determine discrepancies in accuracy associated with ancestry. We hypothesize that RNA-Seq HLA callers are capable of returning high-quality results, but the tools that offer a good balance between accuracy and computational expensiveness for all ancestry groups are yet to be developed. We believe that our study will provide clinicians and researchers with clear guidance to inform their selection of an appropriate HLA caller.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DR2S: An Integrated Algorithm Providing Reference-Grade Haplotype Sequences from Heterozygous Samples 94%
- SOAPTyping: an open-source and cross-platform tool for Sanger sequence-based typing for HLA class I and II alleles 94%
- Frugal alignment-free identification of FLT3-internal tandem duplications with FiLT3r 93%
Similar papers in this journal
- In silico tools for accurate HLA and KIR inference from clinical sequencing data empower immunogenetics on individual-patient and population scales 95%
- An unbiased comparison of immunoglobulin sequence aligners 94%
- Systematic evaluation of transcriptomics-based deconvolution methods and references using thousands of clinical samples 93%
Similar papers in this journal
Similar papers in this journal
- Hypermut 3: Identifying specific mutational patterns in a defined nucleotide context that allows multistate characters 92%
- RNApysoforms: Fast rendering interactive visualization of RNA isoform structure and expression in Python 92%
- Comparative genome analysis using sample-specific string detection in accurate long reads 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.