Discovery of 10,828 new putative human immunoglobulin heavy chain IGHV variants
Martins, F. R.; Pontes, L. A. d. M.; Mendes, T. A. d. O.; Felicori, L. F.
Show abstract
The correct identification of immunoglobulin alleles in genome sequences is a challenge. Nevertheless, it can assist in the study of several human diseases associated with the antibody repertoire and in the development of new therapies using antibody engineering techniques. The advent of next-generation sequencing of human genomes and antibody repertoires enabled the development of several tools for the mapping and identification of new immunoglobulin (Ig) alleles. Some of these tools use 1,000 Genomes (G1K) data for new Ig alleles discovery. However, genome data from G1K present low coverage and variant call problems. Here, a computational screen of immunoglobulin alleles was carried out in the Genome Aggregation Database (gnomAD), the largest high-quality catalogue of variation from 125,748 exomes and 15,708 human genomes. A total of 10,909 putative IGHV alleles were identified, in which 10,828 of them are new and 2,024 appear at least in 6 different alleles from genomes/exomes. The IGHV2-70 was the IGHV gene segment with the largest number of variants described. The majority of the variants were found in the framework 3 and most of them are missense. Interestingly, a large number of variants were found to be population exclusive. A database integrated with a web platform was created (YGL-DB) to store and make accessible the likely new variants found. This available data can help the scientific community to validate new IGHV variants as well as it can shed light on the importance of variants in disease development and immunization protocols.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deciphering of Gorilla gorilla gorilla Immunoglobulin Loci in Multiple Genome Assemblies and Enrichment of IMGT Resources 96%
- Poorly expressed alleles of several human immunoglobulin heavy chain variable (IGHV) genes are common in the human population 95%
- AIRR-C Human IG Reference Sets: curated sets of immunoglobulin heavy and light chain germline genes 95%
Similar papers in this journal
- Understanding SARS-CoV-2 Spike glycoprotein clusters and their impact on immunity of the population from Rio Grande do Norte, Brazil 93%
- Functional prediction and comparative population analysis of variants in genes for proteases and innate immunity related to SARS-CoV-2 infection 93%
- A Sanger-based approach for scaling up screening of SARS-CoV-2 variants of interest and concern 93%
Similar papers in this journal
- CD8 T cell epitope generation toward the continually mutating SARS-CoV-2 spike protein in genetically diverse human population: Implications for disease control and prevention 95%
- A unified classification system for HIV-1 5' long terminal repeats 93%
- Whole Genome Sequencing Analysis of Spike D614G Mutation Reveals Unique SARS-CoV-2 Lineages of B.1.524 and AU.2 in Malaysia 93%
Similar papers in this journal
- KIR2DL4 genetic diversity in a Brazilian population sample: implications for transcription regulation and protein diversity in samples with different ancestry backgrounds 92%
- Mannose-binding lectin gene polymorphisms in the East Siberia and Russian Arctic populations 92%
- Computational identification and characterization of antigenic properties of Rv3899c of Mycobacterium tuberculosis and its interaction with Human leukocyte antigen (HLA) 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.