Back

The Human Canonical Core Histone Catalogue

Susano Pinto, D. M.; Flaus, A.

2019-07-30 molecular biology
10.1101/720235 bioRxiv
Show abstract

Core histone proteins H2A, H2B, H3, and H4 are encoded by a large family of genes distributed across the human genome. Canonical core histones contribute the majority of proteins to bulk chromatin packaging, and are encoded in 4 clusters by 65 coding genes comprising 17 for H2A, 18 for H2B, 15 for H3, and 15 for H4, along with at least 17 total pseudogenes. The canonical core histone genes display coding variation that gives rise to 11 H2A, 15 H2B, 4 H3, and 2 H4 unique protein isoforms. Although histone proteins are highly conserved overall, these isoforms represent a surprising and seldom recognised variation with amino acid identity as low as 77% between canonical histone proteins of the same type. The gene sequence and protein isoform diversity also exceeds commonly used subtype designations such as H2A.1 and H3.1, and exists in parallel with the well-known specialisation of variant histone proteins. RNA sequencing of histone transcripts shows evidence for differential expression of histone genes but the functional significance of this variation has not yet been investigated. To assist understanding of the implications of histone gene and protein diversity we have catalogued the entire human canonical core histone gene and protein complement. In order to organise this information in a robust, accessible, and accurate form, we applied software build automation tools to dynamically generate the canonical core histone repertoire based on current genome annotations and then to organise the information into a manuscript format. Automatically generated values are shown with a light grey background. Alongside recognition of the encoded protein diversity, this has led to multiple corrections to human histone annotations, reflecting the flux of the human genome as it is updated and enriched in reference databases. This dynamic manuscript approach is inspired by the aims of reproducible research and can be readily adapted to other gene families.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
PLOS ONE
5266 papers in training set
Top 7%
23.0%
2
Epigenetics & Chromatin
42 papers in training set
Top 0.1%
19.1%
3
BMC Genomic Data
13 papers in training set
Top 0.1%
6.5%
4
International Journal of Molecular Sciences
494 papers in training set
Top 2%
4.5%
50% of probability mass above
5
BMC Genomics
406 papers in training set
Top 1%
4.5%
6
Biomolecules
100 papers in training set
Top 0.2%
4.2%
7
Genes
144 papers in training set
Top 1.0%
2.7%
8
Biology of Reproduction
36 papers in training set
Top 0.3%
2.2%
9
Journal of Molecular Biology
232 papers in training set
Top 1%
2.2%
10
Cells
249 papers in training set
Top 2%
2.2%
11
PLOS Genetics
862 papers in training set
Top 6%
2.2%
12
Computational and Structural Biotechnology Journal
242 papers in training set
Top 3%
1.8%
13
Mobile DNA
31 papers in training set
Top 0.1%
1.8%
14
Genome Biology
637 papers in training set
Top 6%
1.5%
15
Nucleic Acids Research
1281 papers in training set
Top 9%
1.5%
16
Scientific Reports
3612 papers in training set
Top 63%
1.2%
17
BMC Bioinformatics
457 papers in training set
Top 5%
1.2%
18
Life Science Alliance
285 papers in training set
Top 7%
0.9%
19
Frontiers in Cell and Developmental Biology
233 papers in training set
Top 5%
0.6%
20
Biochimica et Biophysica Acta (BBA) - Molecular Basis of Disease
26 papers in training set
Top 1.0%
0.6%
21
Algal Research
21 papers in training set
Top 0.4%
0.6%
22
Cancers
213 papers in training set
Top 5%
0.6%