The Human Canonical Core Histone Catalogue
Susano Pinto, D. M.; Flaus, A.
Show abstract
Core histone proteins H2A, H2B, H3, and H4 are encoded by a large family of genes distributed across the human genome. Canonical core histones contribute the majority of proteins to bulk chromatin packaging, and are encoded in 4 clusters by 65 coding genes comprising 17 for H2A, 18 for H2B, 15 for H3, and 15 for H4, along with at least 17 total pseudogenes. The canonical core histone genes display coding variation that gives rise to 11 H2A, 15 H2B, 4 H3, and 2 H4 unique protein isoforms. Although histone proteins are highly conserved overall, these isoforms represent a surprising and seldom recognised variation with amino acid identity as low as 77% between canonical histone proteins of the same type. The gene sequence and protein isoform diversity also exceeds commonly used subtype designations such as H2A.1 and H3.1, and exists in parallel with the well-known specialisation of variant histone proteins. RNA sequencing of histone transcripts shows evidence for differential expression of histone genes but the functional significance of this variation has not yet been investigated. To assist understanding of the implications of histone gene and protein diversity we have catalogued the entire human canonical core histone gene and protein complement. In order to organise this information in a robust, accessible, and accurate form, we applied software build automation tools to dynamically generate the canonical core histone repertoire based on current genome annotations and then to organise the information into a manuscript format. Automatically generated values are shown with a light grey background. Alongside recognition of the encoded protein diversity, this has led to multiple corrections to human histone annotations, reflecting the flux of the human genome as it is updated and enriched in reference databases. This dynamic manuscript approach is inspired by the aims of reproducible research and can be readily adapted to other gene families.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Methylation pattern of polymorphically imprinted nc886 is not conserved across mammalia 92%
- THAP11F80L cobalamin disorder-associated mutation reveals normal and pathogenic THAP11 functions in gene expression and cell proliferation 91%
- Progerin Can Induce DNA Damage in the Absence of Global Changes in Replication or Cell Proliferation 91%
Similar papers in this journal
- Exploring the effects of genetic variation on gene regulation in cancer in the context of 3D genome structure 91%
- Predicted Coronavirus Nsp5 Protease Cleavage Sites in the Human Proteome: A Resource for SARS-CoV-2 Research 90%
- Global non-random abundance of short tandem repeats in rodents and primates. 89%
Similar papers in this journal
- Big fish, little fish: N-terminal acetyltransferase Naa40p proteoforms caught in the act 92%
- Unveiling epigenetic regulatory elements associated with breast cancer development 91%
- Cracking the floral quartet code: How do multimers of MIKC C-type MADS-domain transcription factors recognize their target genes? 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.