Back

AmarylOmicBase: A unified transcriptome database to accelerate gene discovery in Amaryllidoideae species

Goncalves dos Santos, K. C.; Merindol, N.; Desgagne-Penix, I.

2025-11-27 bioinformatics
10.1101/2025.11.24.690262 bioRxiv
Show abstract

Amaryllidoideae plants produce structurally diverse and unique alkaloids with potent anti-cholinesterase, antiviral, and antitumor activities, making this subfamily a rich source of pharmaceutical leads. Despite the absence of reference genomes for any Amaryllidoideae species, many enzyme characterization and pathway reconstruction efforts to date have been made possible through transcriptome mining, often requiring bioinformatic expertise and data preprocessing. To facilitate new studies in this subfamily, here we present AmarylOmicBase, a unified transcriptomic database that integrates assemblies, annotations, and expression profiles from 39 studies, covering 29 species across 13 genera of Amaryllidoideae. The AmarylOmicBase includes both published and de novo assemblies generated from published raw data using Trinity or IsoSeq workflows and provides standardized functional annotation and quantitative expression datasets. AmarylOmicBase enables rapid identification of candidate genes, facilitates comparative analyses across species and tissues, and supports pathway reconstruction for specialized metabolism, including Amaryllidaceae alkaloid biosynthesis. By providing ready-to-use datasets and fully reproducible analysis scripts, this resource reduces computational barriers and expands access to transcriptomic information for researchers working on non-model plant species. AmarylOmicBase will accelerate discoveries related to enzyme function, pathway evolution, and the regulatory networks underlying chemical diversity in Amaryllidoideae.

Published in Scientific Data (predicted rank #12) · training set

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.