Genolator: A Multimodal Large Language Model Fusing Natural Language, Genomic, and Structural Tokens for Protein Function Interpretation
Danner, M.; Islam, T.; Begemann, M.; Kraft, F.; Elbracht, M.; Kurth, I.; Krause, J.
Show abstract
BackgroundDecoding the genetic code to unveil its genome functionality is a monumental task which would greatly advance the understanding of disease mechanisms and development of targeted treatment approaches. Although large language models (LLMs) have transformed natural language processing across diverse domains, translating the complex language of DNA into human-readable form remains challenging due to genomic data complexity and unexplored regions of the human genome. Current (genomic) language models either have a solid understanding of natural language or of the genomic code. Models fusing both aspects are largely lacking. ResultsHere we present Genolator, a multimodal large language model that integrates embeddings from DNA sequences, amino acid sequences, and protein structures with natural language queries. Fine-tuned on over 370,000 question-answer pairs generated using abstracted Gene-Ontology (GO) terms, Genolator effectively answers queries regarding protein subcellular localization, molecular function, and biological processes. Evaluation demonstrates high accuracy in confirming or denying protein function associations, outperforming baseline models such as openly available allrounder LLMs like GPT 4.1 as well as smaller domain specific models integrating knowledge from foundation models like Evo2 and ESM-2. Explorations of the Genolators hidden states unveil a biologically and linguistically plausible organization of its learned representations. ConclusionGenolator enhances accessibility to genomic information by enabling natural language interaction with protein data, facilitating biological discovery, and clinical research. It represents a step towards bridging genomic code and human language through the integration of a multimodal LLM.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CaLMPhosKAN: Prediction of General Phosphorylation Sites in Proteins via Fusion of Codon Aware Embeddings with Amino Acid Aware Embeddings and Wavelet-based Kolmogorov Arnold Network 96%
- ProteinBERT: A universal deep-learning model of protein sequence and function 95%
- Embeddings of genomic region sets capture rich biological associations in lower dimensions 95%
Similar papers in this journal
- PLMFit : Benchmarking Transfer Learning with Protein Language Models for Protein Engineering 96%
- Cracking the black box of deep sequence-based protein-protein interaction prediction 95%
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.