NJGPT: A Large Language Model-Driven, User-Friendly Solution for Phylogenetic Tree Construction
Wang, Z.; Huang, H.; Li, T.; Rodrigo, A. G.
Show abstract
MotivationPhylogenetic reconstruction plays an integral part in much of the research in evolutionary biology. Currently, a newly minted phylogeneticist must choose amongst a relatively large array of phylogenetic software, each often with its own analytical routines, inputs and outputs. Our overarching aim is to construct a user-friendly pipeline with the most recent generative AI tool, ChatGPT-4 (at the time of writing) released by OpenAI, that is able to understand queries written in natural language, to build a phylogenetic tree using sequence data. By doing this, we demonstrate how generative AI may be used in phylogenetics, as a proof-of-concept. We also demonstrate the steps needed, presently, to ensure that ChatGPT can build phylogenetic trees accurately. ResultsWe present NJGPT, a phylogenetic tool built using ChatGPT, a Large Language Model (LLM) Generative Pre-trained Transformer (GPT),which employs the Neighbor-Joining method to construct phylogenetic trees. NJGPT simplifies phylogenetic tree construction by allowing users to generate and visualize trees using natural language queries. It supports multiple sequence file formats, matrix calculation models, and gap-deletion methods. To evaluate the performance of NJGPT, we compared output and runtimes with the widely-used phylogenetic software, MEGA. Our results show that NJGPT produces identical trees over a range of sequence lengths and simple models of evolution. However, NJGPT faces visualization issues with datasets over 50 taxa and operational failures with larger datasets due to token limits. NJGPT runtimes were also substantially slower than MEGA; however, NJGPTs user-friendly interface makes it ideal for beginners. AvailabilityThis plugin is available for free at https://chatgpt.com/g/g-1OzP3Qviw-njgpt, the source code is available on GitHub (https://github.com/ZWan622/NJGPT1.0.git) and is implemented using Python Contactzwan622@aucklanduni.ac.nz Supplementary informationSupplementary data are available at Bioinformatics online.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Treerecs: an integrated phylogenetic tool, from sequences to reconciliations 96%
- PoSeiDon: a Nextflow pipeline for the detection of evolutionary recombination events and positive selection 95%
- Build a Better Bootstrap and the RAWR Shall Beat a Random Path to Your Door: Phylogenetic Support Estimation Revisited 95%
Similar papers in this journal
Similar papers in this journal
- Evaluating probabilistic programming and fast variational Bayesian inference in phylogenetics 94%
- DnoisE: Distance denoising by Entropy. An open-source parallelizable alternative for denoising sequence datasets 93%
- Parallel power posterior analyses for fast computation of marginal likelihoods in phylogenetics 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.