Back

CharacTERT: A machine learning tool for classifying hTERT missense variants

Becerra Parra, G.; Pan, Q.; Myung, Y.; Portelli, S.; Nelson, N. E.; Dickinson, J. L.; Lucas, S. E. M.; Holien, J. K.; Bryan, T. M.; Ascher, D. B.

2026-05-20 bioinformatics
10.64898/2026.05.18.725793 bioRxiv
Show abstract

Missense mutations in TERT, the gene encoding the human telomerase catalytic subunit hTERT, are associated with Telomere Biology Disorders (TBDs). Experimentally elucidating the effects of all possible missense variants would be time-consuming and technically challenging. Moreover, current computational predictors are not hTERT-specific and primarily rely on sequence information, failing to capture the complex biological and structural context of the telomerase enzyme. In this work, we developed three machine learning models integrating both sequence- and structure-based features to account for the biological mechanisms of hTERT. Compared to state-of-the-art methods, our best-performing models achieved a higher Matthews Correlation Coefficient of 0.88 on ClinVar and gnomAD curated variants and demonstrated robust sensitivity (0.75) on a dataset curated according to guidelines from the American College of Medical Genetics and Genomics and Association for Molecular Pathology (ACMG/AMP). Feature interpretation highlighted hTERT residue conservation and changes in hydrophobic and weak polar interactions as critical determinants of pathogenicity. Finally, in silico saturation mutagenesis was performed to present a mutational landscape of TERT, available in a user-friendly web server, CharacTERT, which could offer valuable insights into the molecular mechanisms driving TBDs, aid in early diagnosis, as well as guide personalized treatment strategies. CharacTERT is freely available at https://biosig.lab.uq.edu.au/charactert/.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.