The effectiveness of Large Language Models with RAG for auto-annotating phenotype descriptions
Kainer, D.
Show abstract
Ontologies are highly prevalent in biology and medicine and are always evolving. Annotating biological text, such as observed phenotype descriptions, with ontology terms is a challenging and tedious task. The process of annotation requires a contextual understanding of the input text and of the ontological terms available. While text-mining tools are available to assist they are largely based on directly matching words and phrases and so lack understanding of the meaning of the query item and of the ontology term labels. Large Language Models (LLMs), however, excel at tasks that require semantic understanding of input text and therefore may provide an improvement for the auto-annotation of text with ontological terms. Here we describe a series of workflows incorporating OpenAI GPTs capabilities to annotate Arabidopsis thaliana and forest tree phenotypic observations with ontology terms, aiming for results that resemble manually curated annotations. These workflows make use of an LLM to intelligently parse phenotypes into short concepts, followed by finding appropriate ontology terms via embedding vector similarity or via Retrieval-Augmented Generation (RAG). The RAG model is a state-of-the-art approach that augments conversational prompts to the LLM with context-specific data to empower it beyond its pre-trained parameter space. We show that the RAG produces the most accurate automated annotations that are often highly similar or identical to expert-curated annotations. Short descriptionLarge Language Models excel at tasks that require semantic understanding of text. Here we use that capability to auto-annotate plant phenotypes with ontological terms and compare to expert annotation.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 95%
- Linking big biomedical datasets to modular analysis with Portable Encapsulated Projects 94%
- PEPhub: a database, web interface, and API for editing, sharing, and validating biological sample metadata 93%
Similar papers in this journal
- AnnSQL: A Python SQL-based package for fast large-scale single-cell genomics analysis using minimal computational resources 93%
- Enhancing Gene Set Overrepresentation Analysis with Large Language Models 93%
- Gilda: biomedical entity text normalization with machine-learned disambiguation as a service 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.