Back

Bioinformatics Copilot 2.0 for Transcriptomic Data Analysis

Wang, Y.; Zhang, W.; Wong, I.; Lin, S.; Kumar, R.; Wang, A.

2024-08-19 bioinformatics
10.1101/2024.08.15.607673 bioRxiv
Show abstract

Large language models, when integrated into bioinformatics tools, offer valuable support for biologists in the analysis of single-cell transcriptomic data. Here, we introduce Bioinformatics Copilot 2.0, an advanced tool that builds upon its predecessor with four significant enhancements: 1) Internal server processing: users can now process large datasets on an internal server, thereby ensuring data privacy. 2) User-controlled analysis: users can take control of their data analysis when tasks exceed the copilots capabilities.3) Real-time information access: users can access up-to-date information from online resources, such as gene function queries and recent publications, thus enhancing data interpretation. 4) Data analysis and result documentation: the copilot now supports the generation of PDFs that include images and detailed analysis. These advancements represent a step towards building a bioinformatics autopilot. Demonstration videos showcasing these features are available at www.biochemml.com/tools.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.