Back

Deep functional profiling of gene sets using Large Language Models: A blueprint for tailored, context-aware functional annotation

khan, t.; Yurieva, M.; Kabeer, B. S. A.; Toufiq, M.; Rinchai, D.; Palucka, K.; Chaussabel, D.

2024-12-19 bioinformatics
10.1101/2024.12.12.628275 bioRxiv
Show abstract

In this study, we present a proof-of-concept approach leveraging Large Language Models (LLMs) for the deep functional profiling of gene sets, addressing the limitations of traditional functional annotation tools. By employing a stepwise prompting strategy with OpenAIs GPT-4, we systematically retrieve, consolidate, and score immune functions associated with each gene in a module, generating ranked lists of functions with detailed justifications and supporting literature evidence. As a blueprint for this approach, we demonstrate how this in-depth, context-aware analysis can characterize Module M10.4 from the BloodGen3 transcriptional module repertoire, confirming its central role in neutrophil-mediated innate immunity and antimicrobial defense while uncovering a nuanced network of immune functions overlooked by conventional methods. The LLM-based workflow identifies cell types and transcriptional programs driving these functions, offering a more granular and quantitative assessment of the gene sets functional associations compared to direct LLM prompting or pathway enrichment analysis. Our results showcase the potential of LLMs in providing biologically meaningful interpretations of gene expression signatures, accounting for cellular composition changes and transcriptional regulation. This initial proof-of-concept study establishes a foundation for the automated, high-throughput functional profiling of transcriptional signatures. By integrating advanced language models with traditional literature-based validation, we present a powerful strategy for unraveling the complex biology of gene sets and transcriptional modules, ultimately facilitating the translation of systems-level molecular data into actionable knowledge.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.