Back

genomicBERT and data-free deep-learning model evaluation

Chen, T.; Tyagi, N.; Chauhan, S.; Peleg, A. Y.; Tyagi, S.

2023-06-01 bioinformatics
10.1101/2023.05.31.542682 bioRxiv
Show abstract

The genome, which serves as the inherent language directing the blueprint of life, offers significant analysis prospects by combining Natural Language Processing (NLP) and machine learning (ML). Integrating biological sequences with other digital healthcare information has potential to transform data-driven diagnostics. Large language models (LLMs) can be harnessed to decode the genomic language. This endeavor encounters three critical challenges: First, long biomolecular sequences require segmentation into smaller subunits, which is non-trivial since many biological "words" remain unknown. Second, the analysis of extended DNA sequences using LLMs demands a compute-intensive infrastructure. Third, ensuring reproducibility and reusability of modeling workflows remains an unresolved issue. To tackle these challenges, we introduce an empirical DNA tokenisation approach and a versatile, semantic-aware, genome language model --genomicBERT. The model is species-agnostic and operates seamlessly at the DNA or RNA levels. By introducing a reduced and specialized DNA vocabulary, our approach minimizes computational overhead and optimizes performance. Our benchmarking demonstrates that the genomicBERT matches or surpasses the performance of contemporary tools on the same datasets under different experimental conditions. To encourage collaboration and ease of access, we introduce genomicBERT as an integral component of the openly accessible conda package, genomeNLP. Validated across diverse case studies, genomicBERT lowers the barriers to decoding genomic language, relying solely on sequence data to extract meaningful insights. HighlightsO_LIThis novel model offers a compelling solution for DNA sequence analysis by significantly reducing model size and computational costs without compromising performance, setting a new standard for efficient model development. C_LIO_LIWe demonstrate that a powerful vocabulary and tokenization method helps to derive patterns from biological sequence data while accounting for hidden semantic rules. C_LIO_LIOur method is agnostic to species or biomolecule type as it is data-driven. Hence, it can be applied to DNA and RNA C_LIO_LIWe validate the important genomicBERT tokens by mapping back to the biologically significant motifs. C_LIO_LIWe present a publicly available genome language modeling toolkit called genomeNLP, specifically designed to combine computational linguistics and genomics, enabling researchers from biology backgrounds to analyze and interpret genomic sequences effectively. C_LI

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.