Benchmarking freely available human leukocyte antigen typing algorithms across varying genes, coverages and typing resolutions
Thuesen, N. H.; Klausen, M. S.; Gopalakrishnan, S.; Trolle, T.; Renaud, G.
Show abstract
The human leukocyte antigen (HLA) system is a group of genes coding for proteins that are central to the adaptive immune system and identifying the specific HLA allele combination of a patient is relevant in organ donation, risk assessment of autoimmune and infectious diseases and cancer immunotherapy. However, due to the high genetic polymorphism in this region, HLA typing requires specialized methods. We investigated the performance of five next-generation-sequencing (NGS) based HLA typing tools with a non-restricted license namely HLA*LA, Optitype, HISAT-genotype, Kourami and STC-Seq. This evaluation was done for the five HLA loci, HLA-A, -B, -C, -DRB1 and -DQB1 using whole-exome sequencing (WES) samples from 829 individuals. The robustness of the tools to lower coverage was evaluated by subsampling and HLA typing 230 WES samples at coverages ranging from 1X to 100X. The typing accuracy was measured across four typing resolutions. Among these, we present two clinically-relevant typing resolutions, which specifically focus on the peptide binding region. On average, across the five HLA genes, HLA*LA was found to have the highest typing accuracy. For the individual genes, HLA-A, -B and -C, Optitypes typing accuracy was highest and HLA*LA had the highest typing accuracy for HLA-DRB1 and -DQB1. The tools robustness to lower coverage data varied widely and further depended on the specific HLA locus. For all class I loci, Optitype had a typing accuracy above 95% (according to the modification of the amino acids in the functionally relevant portion of the protein) at 50X, but increasing the depth of coverage beyond even 100X could still improve the typing accuracy of HISAT-genotype, Kourami, and STC-seq across all five HLA genes as well as HLA*LAs typing accuracy for HLA-DQB1. HLA typing is also used in studies of ancient DNA (aDNA), which often is based on lower quality sequencing data. Interestingly, we found that Optitypes typing accuracy is not notably impaired by short read length or by DNA damage, which is typical of aDNA, as long as the depth of coverage is sufficiently high.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- SOAPTyping: an open-source and cross-platform tool for Sanger sequence-based typing for HLA class I and II alleles 96%
- MADloy: Robust detection of mosaic loss of chromosome Y from genotype-array-intensity data 94%
- Frugal alignment-free identification of FLT3-internal tandem duplications with FiLT3r 94%
Similar papers in this journal
- In silico tools for accurate HLA and KIR inference from clinical sequencing data empower immunogenetics on individual-patient and population scales 95%
- An unbiased comparison of immunoglobulin sequence aligners 93%
- Hierarchical cell-type identifier accurately distinguishes immune-cell subtypes enabling precise profiling of tissue microenvironment with single-cell RNA-sequencing 93%
Similar papers in this journal
- iCOMIC: a graphical interface-driven bioinformatics pipeline for analyzing cancer omics data 94%
- Kmerator Suite: design of specific k-mer signatures andautomatic metadata discovery in large RNA-Seq datasets. 94%
- IMGT(R) Analysis of the Human IGH Locus: Unveiling Novel Polymorphisms and Copy Number Variations in Genome Assemblies from Diverse Ancestral Backgrounds 93%
Similar papers in this journal
- Accurate multi-population imputation of MICA, MICB, HLA-E, HLA-F and HLA-G alleles from genome SNP data 96%
- Variant calling tool evaluation for variable size indel calling from next generation whole genome and targeted sequencing data 95%
- High-throughput Interpretation of Killer-cell Immunoglobulin-like Receptor Short-read Sequencing Data with PING 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.