Automating Candidate Gene Prioritization with Large Language Models: Development and Benchmarking of an API-Driven Workflow Leveraging GPT-4
khan, T.; Toufiq, M.; Yurieva, M.; Indrawattana, N.; Jittmittraphap, A.; Kosoltanapiwat, N.; Pumirat, P.; Sukphopetch, P.; Vanaporn, M.; Kaber, B.; Palucka, K.; Rinchai, D.; Chaussabel, D.
Show abstract
In this exploratory study, we developed an automated workflow that leverages Large Language Models, specifically GPT-4, to prioritize candidate genes for targeted assay development. The workflow automates interaction with OpenAI models and enables prompt creation, submission. It features customizable prompts designed to evaluate candidate genes based on criteria such as association with biological processes, biomarker potential, and therapeutic implications, which can be tailored for specific diseases or processes. Benchmarking experiments comparing the performance of the Application Programming Interface (API)-based automated prompting approach with manual prompting demonstrated high consistency and reproducibility in gene prioritization results. The automated method exhibited scalability by successfully prioritizing genes relevant to sepsis from the BloodGen3 repertoire, comprising 11,465 genes, distributed among 382 modules. The workflow efficiently identified sepsis-associated genes across the repertoire, revealing distinct gene clusters and providing insights into their distribution within module aggregates and individual modules. This proof-of-concept study demonstrates how LLMs can enhance gene prioritization, streamlining the identification process for targeted assays across various biological contexts. However, it also reveals the need for further validation and highlights the exploratory nature of this work due to scoring inconsistencies and the necessity for manual fact-checking. Despite these challenges, the automated workflow holds promise for accelerating targeted assay development for disease management and paves the way for future research.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- GeneSetCluster 2.0: a comprehensive toolset for summarizing and integrating gene-sets analysis 94%
- GEOlimma: Differential Expression Analysis and Feature Selection Using Pre-Existing Microarray Data 94%
- scMuffin: an R package for disentangling solid tumor heterogeneity from single-cell expression data 93%
Similar papers in this journal
- iCOMIC: a graphical interface-driven bioinformatics pipeline for analyzing cancer omics data 94%
- Predicting Gene Disease Associations With Knowledge Graph Embeddings For Diseases With Curtailed Information 93%
- FLYNC: A Machine Learning-Driven Framework for Discovering Long Non-Coding RNAs in Drosophila melanogaster 93%
Similar papers in this journal
- Changing word meanings in biomedical literature reveal pandemics and new technologies 90%
- A compact encoding of the genome suitable for machine learning prediction of traits and genetic risk scores. 90%
- A Regularized Cox Hierarchical Model for Incorporating Annotation Information in Predictive Omic Studies 90%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.