Back

Gigabyte

GigaScience Press

All preprints, ranked by how well they match Gigabyte's content profile, based on 62 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Long-read sequencing based genomic data of a dipluran species, Occasjapyx japonicus

Asano, T.; Toyoda, A.; Hashimoto, K.; Yokoi, K.

2026-07-21 genomics 10.64898/2026.07.16.716246 medRxiv
Top 0.1%
74.1%
Show abstract

We present the genome dataset of a dipluran species, Occasjapyx japonicus, representing the first dipluran genome assembled using HiFi long-read sequencing technology. The assembled genome is approximately 439.3 Mbp in size, comparable to those of other dipluran species available in public databases. The N50 value of 15.5 Mbp exceeds that reported for other dipluran species. The assembled gene set contains 19,635 genes, a number not significantly different from those estimated in previous analyses of two other dipluran species. Functional gene annotation was conducted using predicted amino acid sequences derived from the gene set. BUSCO analysis indicated that the assembled genome contains the majority of conserved core genes. These findings suggest that the O. japonicus genome and associated data are of sufficient quality to serve as a reference genome. The dataset will be valuable for studies in comparative or evolutionary biology, particularly in understanding hexapod evolution and the emergence of insects.

2
Chromosomal-level genome assembly and single-nucleotide polymorphism sites of black-faced spoonbill Platalea minor

Hong Kong Biodiversity Genomics Consortium, ; Hui, J. H. L.; Chan, T.-F.; Chan, L.; Cheung, S. G.; Cheang, C. C.; Fang, J.; Gaitan-Espitia, J. D.; Lau, S.; Sung, Y. H.; Wong, C.; Yip, K.; Wei, Y.; So, W. L.; Nong, W.; law, s.; Rose-Jeffreys, L.; Crow, P.; Leong, A.; Yip, H.

2024-04-11 genomics 10.1101/2024.04.08.588650 medRxiv
Top 0.1%
72.6%
Show abstract

Platalea minor, the black-faced spoonbill (Threskiornithidae) is a wading bird that is confined to coastal areas in East Asia. Due to habitat destruction, it has been classified by The International Union for Conservation of Nature (IUCN) as globally endangered species. Nevertheless, the lack of its genomic resources hinders our understanding of their biology, diversity, as well as carrying out conservation measures based on genetic information or markers. Here, we report the first chromosomal-level genome assembly of P. minor using a combination of PacBio SMRT and Omni-C scaffolding technologies. The assembled genome (1.24 Gb) contains 95.33% of the sequences anchored to 31 pseudomolecules. The genome assembly also has high sequence continuity with scaffold length N50 = 53 Mb. A total of 18,780 protein-coding genes were predicted, and high BUSCO score completeness (93.7% of BUSCO metazoa_odb10 genes) was also revealed. A total of 6,155,417 bi-allelic SNPs were also revealed from 13 P. minor individuals, accounting for [~]5% of the genome. The resource generated in this study offers the new opportunity for studying the black-faced spoonbill, as well as carrying out conservation measures of this ecologically important spoonbill species.

3
A high-quality de novo genome assembly from a single parasitoid wasp

Ye, X.; Yang, Y.; Tian, Z.; Xu, L.; Yu, K.; Xiao, S.; Yin, C.; Xiong, S.; Fang, Q.; Chen, H.; Li, F.; Ye, G.

2020-07-14 genomics 10.1101/2020.07.13.200725 medRxiv
Top 0.1%
62.3%
Show abstract

Sequencing and assembling a genome with a single individual have several advantages, such as lower heterozygosity and easier sample preparation. However, the amount of genomic DNA of some small sized organisms might not meet the standard DNA input requirement for current sequencing pipelines. Although few studies sequenced a single small insect with about 100 ng DNA as input, it may still be challenging for many small organisms to obtain such amount of DNA from a single individual. Here, we use 20 ng DNA as input, and present a high-quality genome assembly for a single haploid male parasitoid wasp (Habrobracon hebetor) using Nanopore and Illumina. Because of the low input DNA, a whole genome amplification (WGA) method is used before sequencing. The assembled genome size is 131.6 Mb with a contig N50 of 1.63 Mb. A total of 99% Benchmarking Universal Single-Copy Orthologs are detected, suggesting the high level of completeness of the genome assembly. Genome comparison between H. hebetor and its relative Bracon brevicornis shows a high-level genome synteny, indicating the genome of H. hebetor is highly accurate and contiguous. Our study provides an example for de novo assembling a genome from ultra-low input DNA, and will be used for sequencing projects of small sized species and rare samples, haploid genomics as well as population genetics of small sized species.

4
A high-quality de novo genome assembly of Asian Crested Ibis (Nipponia Nippon) using long-read and Hi-C data

yu, y.; Kim, S.-j.; Yoon, C.; Bhak, J.; Kim, C.; Park, H.; Kang, Y.; Kim, Y.; Lee, Y.-j.; Kang, S.-y.; Shin, Y.-u.; Bhak, J.; Jeon, S.

2024-05-02 bioinformatics 10.1101/2024.04.29.591545 medRxiv
Top 0.1%
62.3%
Show abstract

We present TtaoRef1, the highest-quality de novo genome assembly of Asian Crested Ibis (Nipponia Nippon) to date consisting of 134 scaffolds with a length of 1.25 Gb and N50 of 101,183,595 bp. This assembly was generated through the utilization of long-read sequencing and Hi-C data. The assessment of assembly quality, conducted via Benchmarking Universal Single-Copy Orthologs (BUSCO), revealed the presence of 96.8% of completely predicted single-copy genes. TtaoRef1 had 18 times longer N50 value than the previous assembly (ASM70822v1), Furthermore, we conducted the annotation of 24,681 protein-coding genes within the newly assembled genome sequences.

5
A high quality chromosome-level genome assembly for the golden mussel (Limnoperna fortunei)

Rodinho Nunes Ferreira, J. G.; Americo, J. A.; Amaral, D. L. A. S.; Sendim, F.; da Cunha, Y. R.; The Darwin Tree of Life Project Consortium, ; Uliano-Silva, M.; Rebelo, M. d. F.

2022-09-30 genomics 10.1101/2022.09.29.509984 medRxiv
Top 0.1%
60.3%
Show abstract

The golden mussel (Limnoperna fortunei) is a highly adaptive species that causes environmental and socioeconomic losses in invaded areas. Reference genomes have proven to be a valuable resource for studying the biology of invasive species. While the current golden mussel genome has been useful for identifying new genes, its high fragmentation hinders some applications. In this Data Note, we provide the first chromosome-level reference genome for the golden mussel. The genome was built using Hi-C, PacBio HiFi and 10X sequencing data. The final assembly contains 99.4% of its total length assembled to the 15 chromosomes of the species and a scaffold N50 of 97.05 Mb. Approximately 47% of the genome was annotated as repetitive sequences. A total of 34 862 protein-coding genes were predicted, of which 84.7% were functionally annotated. This new high quality genome is expected to support both basic and applied research on this invasive species. Species taxonomyEukaryota; Opisthokonta; Metazoa; Eumetazoa; Bilateria; Protostomia; Spiralia; Lophotrochozoa; Mollusca; Bivalvia; Autobranchia; Pteriomorphia; Mytilida; Mytiloidea; Mytilidae; Arcuatulinae; Limnoperna; Limnoperna fortunei (Dunker, 1857) (NCBI Taxonomy ID: 356393)

6
A Draft Reference Genome for Hirudo verbana, the Medicinal Leech

Paulsen, R. T.; Agany, D. D. M.; Petersen, J.; Davis, C. M.; Ehli, E. A.; Gnimpieba, E.; Burrell, B. D.

2020-12-08 genomics 10.1101/2020.12.08.416024 medRxiv
Top 0.1%
59.2%
Show abstract

The medicinal leech, Hirudo verbana, is a powerful model organism for investigating fundamental neurobehavioral processes. The well-documented arrangement and properties of H. verbanas nervous system allows changes at the level of specific neurons or synapses to be linked to physiological and behavioral phenomena. Juxtaposed to the extensive knowledge of H. verbanas nervous system is a limited, but recently expanding, portfolio of molecular and multi-omics tools. Together, the advancement of genetic databases for H. verbana will complement existing pharmacological and electrophysiological data by affording targeted manipulation and analysis of gene expression in neural pathways of interest. Here, we present the first draft genome assembly for H. verbana, which is approximately 250 Mbp in size and consists of 61,282 contigs. Whole genome sequencing was conducted using an Illumina sequencing platform followed by genome assembly with CLC-Bio Genomics Workbench and subsequent functional annotation. Ultimately, the diversity of organisms for which we have genomic information should parallel the availability of next generation sequencing technologies to widen the comparative approach to understand the involvement and discovery of genes in evolutionarily conserved processes. Results of this work hope to facilitate comparative studies with H. verbana and provide the foundation for future, more complete, genome assemblies of the leech.

7
The genome of the sapphire damselfish Chrysiptera cyanea: a new resource to support further investigation of the evolution of Pomacentrids

Gairin, E.; Miura, S.; Takamiyagi, H.; Herrera, M.; Laudet, V.

2024-11-08 genomics 10.1101/2024.11.06.622371 medRxiv
Top 0.1%
58.7%
Show abstract

The number of high-quality genomes is rapidly growing across taxa. However, it remains limited for coral reef fish of the Pomacentrid family, with most research focused on anemonefish. Here, we present the first assembly for a Pomacentrid of the genus Chrysiptera. Using PacBio long-read sequencing with a coverage of 94.5x, the genome of the Sapphire Devil, Chrysiptera cyanea was assembled and annotated. The final assembly consisted of 896 Mb pairs across 91 contigs, with a BUSCO completeness of 97.6%. 28,173 genes were identified. Comparative analyses with available chromosome-scale assemblies for related species identified contig-chromosome correspondences. This genome will be useful to use as a comparison to study the specific adaptations linked to symbiosis life of the closely related anemonefish. Furthermore, this species is present in most tropical coastal areas in the Indo-West Pacific and could become a model for environmental monitoring. This work will allow to expand coral reef research efforts and highlights the power of long-read assemblies to retrieve high quality genomes.

8
The first genome assembly of the amphibian nematode parasite (Aplectana chamaeleonis)

Han, L.; Liu, T.; He, F.; Hou, Z.

2023-03-23 genomics 10.1101/2023.03.20.533390 medRxiv
Top 0.1%
52.3%
Show abstract

Cosmocercoid nematodes are common parasites of the digestive tract of amphibians. Genomic resources are important for understanding the evolution of a species and the molecular mechanisms of parasite adaptation. So far, no genome resource of Cosmocercoid has been reported. In 2020, a massive Cosmocercoid infection was found in the small intestine of a toad, causing severe intestinal blockage. We morphologically identified this parasite as A. chamaeleonis. Here, we report the first A. chamaeleonis genome with a genome size of 1.04 Gb. The repeat content of this A. chamaeleonis genome is 72.45 %, and the total length is 751 Mb. This resource is fundamental for understanding the evolution of Cosmocercoid and provides the molecular basis for Cosmocercoid infection and control.

9
Chromosome-level genome assembly and annotation of the chaetognath Flaccisagitta enflata

Guiglielmoni, N.; Eitel, M.; Moreau, P.; Krebs, S.; Vermeij, M.; Koszul, R.; Flot, J.-F.

2025-03-17 genomics 10.1101/2025.03.14.643333 medRxiv
Top 0.1%
50.7%
Show abstract

Chaetognaths, a phylum of enigmatic marine predators, present a significant challenge to phylogenetic reconstruction due to their uncertain evolutionary placement. While transcriptome analyses have suggested affinities with the Gnathifera clade, genomic data for this group remain scarce, hindering a comprehensive understanding of their evolution. Here, we present the first chromosome-level genome assembly of Flaccisagitta enflata, a species within the Aphragmophora order. The genome assembly includes 9 chromosome candidates with a total size of 794 Mb and a BUSCO score of 91.3% against the Metazoa lineage. This high-quality genome assembly provides a crucial resource for comparative genomic analyses within Chaetognatha and the broader Gnathifera clade, and it will facilitate investigations into chaetognath evolution and their phylogenetic relationships, addressing long-standing questions regarding their placement within the animal kingdom.

10
Highly contiguous reference genome assembly of the endangered Orces blue whiptail Holcosus orcesi

Pozo, G.; Cisneros-Heredia, D. F.; Barragan-Orbe, D.; Sanchez-Nivicela, J. C.; Arbelaez, E.; Torres, M.

2026-05-16 genomics 10.64898/2026.05.14.725226 medRxiv
Top 0.1%
46.0%
Show abstract

Holcosus orcesi, the Orces Blue Whiptail, is a Critically Endangered lizard endemic to the upper Jubones River basin in southern Ecuador. Restricted to a narrow elevational range within semi-arid Andean shrublands, it represents one of the few montane members of a predominantly lowland lineage. Here we present the first high-quality reference genome for H. orcesi, generated using Oxford Nanopore Technologies long-read sequencing. The assembly spans 1.68 Gb across only 91 contigs, with an N50 of 76.2 Mb and a BUSCO completeness of 96.8%, making it among the most contiguous and complete squamate genomes to date. Structural annotation predicted 25,682 genes, of which 85% showed homology to known proteins and 45% were assigned Gene Ontology terms. Repetitive elements accounted for 46.3% of the genome, with LINEs representing the predominant class. This genome provides a foundational resource for future evolutionary, comparative and conservation-genomic research of H. orcesi and other mountain reptiles, enabling studies of population genomics, local adaptation, and genomic erosion in isolated populations. By expanding the genomic representation of tropical montane reptiles, this work helps address longstanding phylogenetic and geographic gaps in global biodiversity genomics and provides a foundation for evidence-based conservation of H. orcesi and related taxa.

11
A chromosome-level reference genome for Pacific herring (Clupea pallasii) from the Bering Sea

Timm, L. E.; Hsieh, Y.; Lopez, J. A.; Almgren, S. A.; Glass, J. R.

2026-02-10 genomics 10.64898/2026.02.09.704930 medRxiv
Top 0.1%
45.6%
Show abstract

Pacific herring (Clupea pallasii) serve as a critical trophic link between plankton and many marine species targeted by fisheries. With a broad distribution throughout the North Pacific Ocean, from the Arctic to temperate latitudes, herring hold ecological, economic, and cultural importance. Despite this importance, genomic resources for this species, such as reference genome sequences, have only recently become available. To date, only one scaffold-level reference genome, representing a specimen from the Gulf of Alaska (Vancouver; 1,379 scaffolds), has been published to NCBI. Addressing this data gap, we produced a high quality 795Mb genome sequence organized into 26 chromosomes combining long read sequencing with short read sequencing of proximity ligation libraries. Our assembly is highly complete (BUSCO score of 97.7%) and contiguous (922 contigs, N50 = 7,338,470, L50 = 38; 26 scaffolds, N50 = 31,494,017; L50 = 12). Pacific herring south of the Aleutian Islands and the Alaska Peninsula are genetically differentiated from those in the Bering Sea, making a reference genome from the eastern Bering Sea an important addition to the Pacific herrings genomic toolbox.

12
Improvements to the Gulf Pipefish Syngnathus scovelli Genome

Ramesh, B.; Small, C.; Healey, H.; Johnson, B.; Barker, E.; Currey, M.; Bassham, S.; Myers, M.; Cresko, W.; Jones, A.

2023-01-24 genomics 10.1101/2023.01.23.525209 medRxiv
Top 0.1%
44.2%
Show abstract

The Gulf pipefish Syngnathus scovelli has emerged as an important species in the study of sexual selection, development, and physiology, among other topics. The fish family Syngnathidae, which includes pipefishes, seahorses, and seadragons, has become an increasingly attractive target for comparative research in ecological and evolutionary genomics. These endeavors depend on having a high-quality genome assembly and annotation. However, the first version of the S. scovelli genome assembly was generated by short-read sequencing and annotated using a small set of RNA-sequence data, resulting in limited contiguity and a relatively poor annotation. Here, we present an improved genome assembly and an enhanced annotation, resulting in a new official gene set for S. scovelli. By using PacBio long-read high-fidelity (Hi-Fi) sequences and a proximity ligation (Hi-C) library, we fill small gaps and join the contigs to obtain 22 chromosome-level scaffolds. Compared to the previously published genome, the gaps in our novel genome assembly are smaller, the N75 is much larger (13.3 Mb), and this new genome is around 95% BUSCO complete. The precision of the gene models in the NCBIs eukaryotic annotation pipeline was enhanced by using a large body of RNA-Seq reads from different tissue types, leading to the discovery of 28,162 genes, of which 8,061 were non-coding genes. This new genome assembly and the annotation are tagged as a RefSeq genome by NCBI and thus provide substantially enhanced genomic resources for future research involving S. scovelli.

13
Closely related, yet phenotypically different - Genome assemblies of two sister species of widow spiders: Latrodectus hasselti and L. katipo, Theridiidae

Ivanov, V.; Uludag, K. O.; Schöneberg, Y.; Schneider, J. M.; Kennedy, S.; Hamadou, A. B.; Vink, C. J.; Krehenwinkel, H.

2026-04-21 genomics 10.64898/2026.04.17.719154 medRxiv
Top 0.1%
44.1%
Show abstract

Widow spiders of the genus Latrodectus are important animals for biomedical, pest and conservation research. Here, we present the assembled genomes of two closely related Latrodectus species: the Australian L. hasselti and the New Zealand endemic L. katipo. The genome of L. katipo consists of 13 scaffolds likely corresponding to chromosomes (90% of the total length) and 1267 short scaffolds (10%). It has a total length of 1.5 Gbp and BUSCO of 94.9%. The genome of L. hasselti consists of 379 scaffolds and has a total length of 1.7 Gbp and a BUSCO score of 95.4%. The repeat content is very similar in both genomes with a total proportion of 37.2% for L. katipo and 39.9% for L. hasselti. Genome annotation predicted 12706 and 15111 genes for L. katipo and L. hasselti respectively. An ortholog analysis shows large overlap between orthogroups suggesting either duplication events in L. hasselti or loss of genes in L. katipo.

14
Characterization of Wnt Signaling Genes in Diaphorina citri

Vosburg, C.; Reynolds, M.; Noel, R.; Shippy, T.; Hosmani, P. S.; Flores-Gonzalez, M.; Mueller, L. A.; Hunter, W. B.; Brown, S. J.; D'elia, T.; Saha, S.

2020-09-22 genomics 10.1101/2020.09.21.306100 medRxiv
Top 0.1%
40.9%
Show abstract

The Asian citrus psyllid, Diaphorina citri, is an insect vector that transmits Candidatus Liberibacter asiaticus, the causal agent of the Huanglongbing (HLB) or citrus greening disease. This disease has devastated Floridas citrus industry and threatens Californias industry as well as other citrus producing regions around the world. To find novel solutions to the disease, a better understanding of the vector is needed. The D. citri genome has been used to identify and characterize genes involved in Wnt signaling pathways. Wnt signaling is utilized for many important biological processes in metazoans, such as patterning and tissue generation. Curation based on RNA sequencing data and sequence homology confirm twenty four Wnt signaling genes within the D. citri genome, including homologs for beta-catenin, Frizzled receptors, and seven Wnt-ligands. Through phylogenetic analysis, we classify D. citri Wnt-ligands as Wg/Wnt1, Wnt5, Wnt6, Wnt7, Wnt10, Wnt11, and WntA. The D. citri version 3.0 genome with chromosomal length scaffolds reveals a conserved Wnt1-Wnt6-Wnt10 gene cluster with gene configuration similar to that in Drosophila melanogaster. These findings provide a greater insight into the evolutionary history of D. citri and Wnt signaling in this important hemipteran vector. Manual annotation was essential for identifying high quality gene models. These gene models can further be used to develop molecular systems, such as CRISPR and RNAi, that target and control D. citri populations, to manage the spread of HLB. Manual annotation of Wnt signaling pathways was done as part of a collaborative community annotation project (https://citrusgreening.org/annotation/index).

15
Chromosome-level Genome Assembly of the South African Lion (Panthera leo melanochaita)

Hadebe, S.; Tshilate, T. S.; Hlongwane, N.; Nesengani, L. T.; Mdyogolo, S.; Molotsi, A.; Smith, R. M.; Labuschagne, K.; Masebe, T.; Mapholi, N.

2026-03-12 genomics 10.64898/2026.03.10.710750 medRxiv
Top 0.1%
40.0%
Show abstract

The lion (Panthera leo melanochaita) is one of the most iconic species and part of the big five, maintaining ecological balance and a major wildlife-based tourism attraction in South Africa. Despite its importance, it is currently threatened by rapid population decline and increasing population fragmentation. Therefore, there is a need for a high-quality genomic resource that captures the diverse genetic landscape of South African lion populations. To address this, we present a high-quality genome assembly of the lion, generated using PacBio HiFi and Omni-C sequencing technologies. The final assembly comprises 2.45 Gb, with a scaffold N50 of 148 Mb and a contig N50 of 22 Mb. Remarkably, 94.8% of the genome is anchored to 19 scaffolds, reflecting the high degree of contiguity and near-complete chromosomal level. Completeness assessment of the genome showed 98.2% BUSCO, 98.2% k-mer completeness and QV of 65.6 underscoring high accuracy and biological integrity. Genome annotation predicted 831.4 Mb (33.9%) of repetitive sequences and 21 739 protein-coding genes. This work provides high-quality genomic resource to establish a foundation for future population genomic and conservation-focused investigations of the lion populations in South Africa.

16
Systematic functional annotation workflow for insects

Bono, H.; Sakamoto, T.; Kasukawa, T.; Tabunoki, H.

2022-06-11 bioinformatics 10.1101/2022.05.12.490705 medRxiv
Top 0.1%
39.7%
Show abstract

Next generation sequencing has revolutionized entomological study, rendering it possible to analyze the genomes and transcriptomes of non-model insects. However, use of this technology is often limited to obtaining nucleotide sequences of target or related genes, with many of the acquired sequences remaining unused because other available sequences are not sufficiently annotated. To address this issue, we have developed a functional annotation workflow for transcriptome-sequenced insects to determine transcript descriptions, which represents a significant improvement over the previous method (functional annotation pipeline for insects). The developed workflow attempts to annotate not only the protein sequences obtained from transcriptome analysis but also the ncRNA sequences obtained simultaneously. In addition, the workflow integrates the expression level information obtained from transcriptome sequencing for application as functional annotation information. Using the workflow, functional annotation was performed on the sequences obtained from transcriptome sequencing of stick insect (Entoria okinawaensis) and silkworm (Bombyx mori), yielding richer functional annotation information than that obtained in our previous study. The improved workflow allows more comprehensive exploitation of transcriptome data and is applicable to other insects because the workflow has been openly developed on GitHub. Simple SummaryThe function of all genes encoded in the genome should be studied for genome editing. The genome editing technology can speeds up insect research for functional analysis of genes. Our knowledge about the functional information of genes is still incomplete currently while genome sequencing of an organism can be completed. The functional information has been annotated based solely on the information that has been obtained from the result of previous biological research. However, this information will be important in determining the target genes for genome editing. In particular, it is very important that this information is in machine-readable form because computer programs mainly parse this information for the understanding of biological systems. In this paper, we describe a workflow-based method for annotating gene functions in insects that make use of transcribed sequence information as well as reference genome and protein sequence databases. Using the developed workflow, we annotated functional information of Japanese stick insect and silkworm, including gene expression as well as sequence analysis. The functional annotation information obtained by the workflow will greatly expand the possibilities of entomological research using genome editing.

17
Improved chromosome level genome assembly of the Glanville fritillary butterfly (Melitaea cinxia) based on SMRT Sequencing and linkage map.

Blande, D.; Smolander, O.-P.; Ahola, V.; Rastas, P.; Tanskanen, J.; Kammonen, J. I.; Oostra, V.; Pellegrini, L.; Ikonen, S.; Dallas, T.; DiLeo, M. F.; Duplouy, A.; Duru, I. C.; Halimaa, P.; Kahilainen, A.; Kuwar, S. S.; Karenlampi, S. O.; Lafuente, E.; Luo, S.; Makkonen, J.; Nair, A.; Celorio-Mancera, M. d. l. P.; Pennanen, V.; Ruokolainen, A.; Sundell, T.; Tervahauta, A. I.; Twort, V.; van Bergen, E.; Osterman-Udd, J.; Paulin, L.; Frilander, M. J.; Auvinen, P.; Saastamoinen, M.

2020-11-04 genomics 10.1101/2020.11.03.364950 medRxiv
Top 0.1%
39.2%
Show abstract

The Glanville fritillary (Melitaea cinxia) butterfly is a long-term model system for metapopulation dynamics research in fragmented landscapes. Here, we provide a chromosome level assembly of the butterflys genome produced from Pacific Biosciences sequencing of a pool of males, combined with a linkage map from population crosses. The final assembly size of 484 Mb is an increase of 94 Mb on the previously published genome. Estimation of the completeness of the genome with BUSCO, indicates that the genome contains 93 - 95% of the BUSCO genes in complete and single copies. We predicted 14,830 gene models using the MAKER pipeline and manually curated 1,232 of these gene models. The genome and its annotated gene models are a valuable resource for future comparative genomics, molecular biology, transcriptome and genetics studies on this species.

18
The reference genome of the paradise fish (Macropodus opercularis)

Fodor, E.; Okendo, J.; Szabo, N.; Szabo, K.; Czimer, D.; Tarjan-Racz, A.; Szeverenyi, I.; Low, B. W.; Liew, J. H.; Koren, S.; Rhie, A.; Orban, L.; Miklosi, A.; Varga, M.; Burgess, S. M.

2023-08-10 genomics 10.1101/2023.08.10.552018 medRxiv
Top 0.1%
39.2%
Show abstract

Over the decades, a small number of model species, each representative of a larger taxa, have dominated the field of biological research. Amongst fishes, zebrafish (Danio rerio) has gained popularity over most other species and while their value as a model is well documented, their usefulness is limited in certain fields of research such as behavior. By embracing other, less conventional experimental organisms, opportunities arise to gain broader insights into evolution and development, as well as studying behavioral aspects not available in current popular model systems. The anabantoid paradise fish (Macropodus opercularis), an "air-breather" species from Southeast Asia, has a highly complex behavioral repertoire and has been the subject of many ethological investigations, but lacks genomic resources. Here we report the reference genome assembly of Macropodus opercularis using long-read sequences at 150-fold coverage. The final assembly consisted of {approx}483 Mb on 152 contigs. Within the assembled genome we identified and annotated 20,157 protein coding genes and assigned {approx}90% of them to orthogroups. Completeness analysis showed that 98.5% of the Actinopterygii core gene set (ODB10) was present as a complete ortholog in our reference genome with a further 1.2 % being present in a fragmented form. Additionally, we cloned multiple genes important during early development and using newly developed in situ hybridization protocols, we showed that they have conserved expression patterns.

19
High-quality genome assembly of a white eared pheasant individual and related functional and genetics data resources

Wu, S.; Wang, K.; Dou, T.; Yuan, S.; Wu, D.-D.; Su, Z.; Ge, C.; Jia, J.

2023-11-13 bioinformatics 10.1101/2023.11.09.566452 medRxiv
Top 0.1%
38.8%
Show abstract

White eared pheasant (WT), (Crossoptilon crossoptilon), inhibiting at high altitudes (3000[~]4,300 m), is a Galliformes bird native to the Qinghai, Sichuan, Yunnan and Tibet Province of China. Due to the difficulty of sequencing the precious species, there is no high-quality genome assembly for the species, hampering the understanding of their genetic mechanisms. To fill the gap, we sequenced and assembled a WT individual using Illumina short reads, PacBio long reads and Hi-C reads. With a contig N50 of 19.63 Mb, scaffold N50 of 29.59 Mb, total length of 1.02 Gb and BUSCO completeness of 97.2%, the assembly is highly complete. Evaluation shows that the assembly is at chromosome-level with only six gaps. Thus, our assembly provides a valuable genetic resource for the crossoptilon species. To further provide resources for gene annotation and population genetics analysis, we also sequenced transcriptomes of 20 tissues of the WT individual and re-sequenced another 10 individuals of WT. Our assembled WT genome and the sequencing data can be valuable resources to study the crossoptilon species.

20
A Highly Contiguous Reference Genome for Scalesia gordilloi (Asteraceae), a Critically Endangered Plant Endemic to the Galapagos Islands

Pozo, G.; Rivas-Torres, G.; Velez-Darquea, E.; Barragan-Orbe, D.; Torres, M. d. L.

2026-06-29 genomics 10.64898/2026.06.25.734018 medRxiv
Top 0.1%
38.8%
Show abstract

Scalesia gordilloi is a critically endangered species endemic to San Cristobal Island in the Galapagos archipelago and represents one of the most unique and vulnerable lineages within the adaptive radiation of the genus Scalesia. Despite its evolutionary distinctiveness and conservation importance, no genomic resources have been available for this species. Here, we present the first high-quality reference genome of S. gordilloi, generated using Oxford Nanopore long-read sequencing. Across three PromethION R10.4.1 flow cells, we obtained 80.5 Gb of long reads (~25X coverage), which enabled a highly contiguous 3.61 Gb assembly composed of only 549 contigs and an N50 of 106.6 Mb. BUSCO completeness reached 98.6%, with assembly metrics comparable to other high-quality Asteraceae genomes. Repeat annotation revealed that 76.2% of the genome is composed of interspersed elements, dominated by LTR retrotransposons. Structural annotation resulted in 47,913 high-confidence protein-coding genes, consistent with expectations for large, repetitive Asteraceae genomes. This genome provides a critical foundation for conservation genomics, enabling assessments of genetic diversity, inbreeding, and adaptive potential in the species. It further establishes a framework for comparative genomics across the Scalesia radiation and supports future efforts to protect and restore one of the most threatened plant lineages of the Galapagos Islands.