Gigabyte
● GigaScience Press
Preprints posted in the last 30 days, ranked by how well they match Gigabyte's content profile, based on 62 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit.
Backlund, A. E.; Nielsen, J.; Pulford, J.; Cook, B.; Anderson, J.; Robert, M.; Thompson, J. S.; Rele, C. P.; Wittke-Thompson, J. K.
Show abstract
Gene model for the ortholog of raptor in the May 2011 (Agencourt Dere_CAF1/DereCAF1) Genome Assembly (GenBank Accession: GCA_000005135.1) of Drosophila erecta. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.
Lieser, B. C.; Lose, B.; Kiser, C. A.; Butterfield, S.; Laschober, L.; Laskowski, L. F.; Nielsen, J.; Pulford, J.; Thompson, J. S.; Rele, C. P.; Wittke-Thompson, J. K.
Show abstract
Gene model for the ortholog of raptor in the D. grimshawi May 2011 (Agencourt dgri_caf1/DgriCAF1) Genome Assembly (GenBank Accession: GCA_000005155.1) of Drosophila grimshawi. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.
Aguiar, A. P.
Show abstract
The preparation of multi panel figures remains a labor intensive step in scientific publication. Albeit there are specific tools available to solve this problem, they are often highly specialized, difficult to install, or time consuming to learn. Griphus is a standalone graphical application designed for rapid composition and experimentation with multi panel figures, developed by and for zoological taxonomists. Functions specifically designed for multi panel composition include automatic figure numbering and placement, aspect ratio operations, spacers, layout rotation, layout suggestions, and automatic generation of figure legends, including scale bar descriptions. The software can perform both spatial interpretation of images on the canvas and work with a simple, editable layout formula. It also enables instant multi panel composition, with numbered images and automatic contrast selection for the numbers, obtained simply by loading images. User defined parameters such as target printable dimensions, resolution, spacing, and color mode are preserved throughout the work. The program produces coordinated outputs consisting of the final composite figure, a readable file describing the layout structure, and a .gri file storing images, transformations, and parameters for exact regeneration. Griphus is intended as a complementary tool to professional image software, providing a simple and efficient environment for constructing high quality multi panel figures.
Perez, J.; Giunta, A. A.; Wittke-Thompson, J. K.
Show abstract
Gene model for the ortholog of tango (tgo) in the Sep. 2015 (UC Berkeley ASM127793v1/DbusGB1) Genome Assembly (GenBank Accession: GCA_001277935.1) of Drosophila busckii. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.
Lawson, M. E.; Sanow, K. A.; Martinand, I.; Fratian, M.; Matura, M.; Rele, C. P.; Reed, L. K.; Thompson, J. S.; O'Rourke, K. S.
Show abstract
Gene model for the ortholog of Density regulated protein (DENR) in the Apr. 2013 (BCM-HGSC/Deug_2.0) (DeugGB2) Genome Assembly (GenBank Accession: GCA_000236325.2) of D. eugracilis. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.
Tran Nguyen, A. H.; Ha, G.-H.; Tran, D.-P.; Le, N. T.; Glendining, S.; Fitzgibbon, Q.; Herzig, V.; Luu, P.-L.; Ventura, T.
Show abstract
The slipper lobster (Thenus australiensis) is rapidly emerging as a high-potential species for commercial aquaculture. Because females exhibit superior growth characteristics due to less frequent moulting after sexual maturity, developing monosex breeding strategies is highly desirable for industry profitability. However, the lack of genomic resources and early sex-identification tools has hindered this development. Here, we report the first draft male genome assembly for T. australiensis, generated using a combination of whole-genome shotgun sequencing, DArT-seq, and multi-tissue transcriptomics. The curated assembly spans 0.913 Gbp with high functional completeness (93.0% BUSCO), providing a robust repertoire of 30,100 protein-coding genes. Through k-mer subtraction and population-level DArT-seq genotyping, we provide definitive evidence that T. australiensis utilizes an XX/XY sex-determination system. Crucially, by identifying male-specific structural variations within a neo-Y locus, we developed a diagnostic PCR assay targeting a male-exclusive sequence. This 171 bp marker achieved 100% accuracy in phenotypic sex identification across wild-caught populations. Ultimately, these foundational genomic resources, combined with a highly reliable molecular sexing tool, provide the critical framework necessary for early sex sorting, broodstock management, and the commercial advancement of monosex slipper lobster farming.
Turner, T. L.
Show abstract
This study presents a systematic revision of the suborder Astrophorina for the temperate Pacific coast of the United States and Canada. Major findings include a reduction in the number of species previously thought to range into the region from Japan; validation of most Geodia species erected by Lendenfeld (1910), which were later synonymized by de Laubenfels (1932); the formal description of 10 new species (Poecillastra alaskensis sp. nov., Vulcanella explorata sp. nov., Vulcanella rupta sp. nov., Stelletta cardenasi sp. nov., Stelletta nicolenya sp. nov., Stelletta limuwensis sp. nov., Dercitus (Stoeba) giveni sp. nov., Penares anyapax sp. nov., Penares foxi sp. nov., and Thenea diastra sp. nov.); and one new combination, Penares orientalis comb. nov. Extensive SCUBA-based collection efforts yielded new samples for 11 of the 26 species identified in the region, which enabled an integrative taxonomic approach that combined field photography, fresh material for DNA sequencing, and improved characterization of species ranges and morphological variability in previously described taxa. Illumina sequencing generated complete nuclear ribosomal haplotypes for five species, while Sanger sequencing of the 28S and cox1 loci placed 20 of the 26 species within molecular phylogenies. The use of very short "mini-barcode" amplicons also enabled sequence recovery from historic type specimens up to 137 years old. This study additionally reports the discovery of sponge grounds of abundant, large Geodia at diving depths in Southern California. Together, these results substantially advance our understanding of global astrophorid diversity and systematics, and the biogeography of sponge diversity in the Northeast Pacific. Note about species names: this pre-print is not intended to be a publication of the associated species names for the purposes of zoological nomenclature.
Horikawa, K.; Savkin, K.; Rower, L.; Hodge, L.; Warren, T. L.
Show abstract
Long-distance movement in insects has crucial impacts on agriculture, human health, and biodiversity. Although it was long assumed that only large, specialist insects had the navigation capacity to support long-distance dispersal, recent studies have demonstrated that smaller insects, such as the tiny fruit fly Drosophila melanogaster, can maintain extended, straight paths while flying or walking. This raises the question of whether other Drosophila species possess the navigation capacity to support extended dispersal. Resolving this question is particularly important for Drosophila suzukii(spotted-wing drosophila), a potent pest species that causes enormous damage worldwide to ripe fruit and berries. Spotted-wing drosophila has been thought to lack a capacity for long-distance dispersal, as prior studies have estimated maximal daily dispersal distances of less than 90 m. We developed a system to continuously track the flight trajectories of magnetically tethered D. suzukii relative to a discrete, overhead LED that mimicked the sun. We found that flies maintained remarkably straight flight headings that varied unpredictably across individuals. Male and female D. suzukii exhibited a similar navigation capacity; both sexes responded to rotation of a discrete sun stimulus with compensatory turns to maintain a stable relative heading. Our results suggest that D. suzukiihas an underappreciated capacity for rapid, radial dispersal, which could exceed 250 m in 15 min. This capacity may contribute to the pest species' invasiveness and its reliable, annual re-establishment in seasonally intolerable climates. Our findings highlight the importance of developing area-wide, regional strategies to manage the impacts of D. suzukii.
Torresen, O. K.; Mysterud, A.; Skage, M.; Danneels, B.; Strand, M. A.; Ferrari, G.; Tooming-Klunderud, A.; Jakobsen, K. S.
Show abstract
We describe a chromosome-level, haplotype-resolved genome assembly from a male European moose (Alces alces alces). The assembly comprises two pseudo-haplotypes of 3,148 Mb and 3,112 Mb, with sex chromosomes in haplotype one, and 33 autosomes in each haplotype (68 in total). Assembly completeness is high (BUSCO 98.3% and 95.7%), with 21,496 and 20,498 annotated protein-coding genes for haplotypes one and two, respectively. This genome assembly is the most complete so far generated for European moose.
Molligan, J.; Sylvestre, F.; Perez-Lopez, E.
Show abstract
The potato leafhopper, Empoasca fabae (Harris, 1841), is a highly polyphagous, migratory insect pest of eastern North America that feeds on more than 200 herbaceous and woody plant species, causing substantial losses to forage and field crops. Despite its agricultural and ecological importance, no genome has been available for this species. Here, we present the first chromosome-level genome assembly of E. fabae, generated from Oxford Nanopore long reads, Illumina short reads, and Omni-C proximity-ligation data. The final assembly spans 908 Mb across 132 scaffolds, with 99.8% of the assembly captured in ten chromosome-length scaffolds (nine autosomes and an X chromosome) with a scaffold N50 of 96.2 Mb. The assembly is highly complete, recovering 92.4% of conserved hemipteran single-copy orthologs, and is composed of 47.6% repetitive sequence, dominated by long terminal repeat retrotransposons and unclassified elements. Read-depth comparison between male and female individuals supports assignment of a single sex-linked chromosome, consistent with an XO sex-determination system. BRAKER3 gene annotation predicted 31,406 protein-coding genes after retaining the longest isoform per locus. Comparative genome analysis against the two closest related Typhlocybinae species with genomes available, Matsumurasca onukii and Hebata decipiens, revealed extensive chromosome-scale collinearity, while defining a shared core gene repertoire. This reference genome provides a foundation for comparative and population genomic studies and for investigating genetic traits in this economically important crop pest species.
Ellerstrand, S. J.; Churcher, A. M. J.; Kutschera, V. E.; Hansson, B.
Show abstract
Sex chromosomes are central to many ecological and evolutionary processes. Evidence has accumulated that sex chromosome systems vary extensively in age, turnover and transitions, motivating renewed efforts to study the diversity of sex chromosome systems across the tree of life. However, successful genomic detection of sex chromosomes depends on several factors, including the size and divergence time, background genetic diversity, and the number of sequenced females and males. In addition, technical challenges associated with sequencing and analysing the sex-limited Y/W chromosome remain. Here, we present PhaseWY, an automated Snakemake pipeline that uses whole-genome sequencing data from multiple female and male individuals to identify sex-chromosomal regions and extract the corresponding Y/W sequences. PhaseWY (i) detects sex differences in alignment depth, (ii) applies read-based and statistical haplotype phasing, (iii) identifies sex-linked regions using haplotype clustering, and (iv) subsets autosomal, X/Z- and Y/W-linked variants for downstream analyses. We applied PhaseWY to simulated data to benchmark factors influencing sex-linkage detection and successful extraction of Y/W-linked variants. To demonstrate its practical utility, we further applied PhaseWY to the neo-sex chromosome system in Alauda larks (Alaudidae) and performed a range of downstream analyses demonstrating the scope of applications of the PhaseWY output. We conclude that PhaseWY provides an easy-to-use and reproducible tool for population-genomic analyses in non-model organisms, with particular importance for advancing our understanding of sex-chromosome evolution.
Cascini, M.; Simpson, L.; Worboys, S.; Worboys, W.; Guja, L.; Knapp, Z.; Bredell, P.; Percival, J.; Rossetto, M.; Crayn, D.
Show abstract
A core aim of ex situ conservation is to represent wild genetic diversity in managed living collections. For the climate-threatened tropical montane cloud forest (TMCF) flora of northeast Australia, an ex situ metacollection of plants and seeds has been established by the Tropical Mountain Plant Science (TroMPS) project. In this study we used reduced-representation sequencing (DArTseq) of wild, herbarium, and ex situ material alongside provenance information for ten species, to pursue two central aims: to characterise landscape-scale genetic structure across species' ranges, and to evaluate how well the assembled metacollections represent that wild diversity. Analyses revealed consistent patterns of genetic differentiation among mountain top populations across multiple species, reflecting the isolating influence of lowland gaps between upland habitats, with the degree of differentiation varying among species. These results provide the first genetic baseline for Australian TMCF flora and reinforce the importance of treating individual mountain top populations as distinct units for conservation management. Additionally, the project provided valuable insights into the logistical challenges of coordinated multi-institutional collecting, informing strategies for metacollection design more broadly. Evaluation of the metacollection revealed both strengths and gaps in representation across species, providing an evidence base to refine the current holdings and guide future targeted collecting to strengthen their long-term conservation value.
Mei, C.; Ness, J.; Nakai, K.; Wunderlich, Z.
Show abstract
Developmental processes depend on carefully coordinated gene expression. Expression is modulated by the binding of transcription factors (TFs) to cis-regulatory elements (CREs), like enhancers and promoters. Many computational and experimental approaches have been developed to find CREs, particularly enhancers, in the genome, each with strengths and caveats. Given the increasing availability of ATAC-seq data and methods to find TF binding therein, we hypothesized that we could use TF footprinting tools to find clusters of TF binding events within accessible chromatin that may act as CREs. Using Drosophila anterior-posterior patterning network as a test bed, we used a digital genomic footprinting tool (DGT), TOBIAS, on previously published early embryo ATAC-seq data to characterize the TF footprint landscape of 16 TFs essential for embryonic patterning. Even in this system, with its extensive enhancer annotation, most footprinted TF binding sites lie outside of known enhancers, with intergenic and intronic regions hosting the highest TF footprint count, albeit at low density. To find potential novel enhancers, we identified high-density TF footprint clusters that are highly conserved and overlap with active enhancer histone mark signals. Five high confidence candidates were selected for reporter assay validation and all five were found to drive spatially patterned expression in the embryo. This study shows that even in a highly characterized system, the analysis of footprinted TF binding sites in ATAC-seq data can uncover new regulatory regions and suggests this approach may be helpful in using existing ATAC-seq data to find novel CREs. ARTICLE SUMMARYGiven the increasing availability of ATAC-seq datasets, workflows to exploit the data to uncover new cis-regulatory elements (CREs), including enhancers, are valuable. Using early anterior-posterior patterning in the Drosophila embryo as a test case, we find that previously published transcription factor footprinting tools and ATAC-seq data can be analyzed to yield new candidate CREs. Experimental validation confirms the activity of selected candidate CREs, suggesting that existing data can be analyzed to find novel regulatory elements.
TOUCEDO, R.; Zhu, Y.; Moledo, S.; Gambon Deza, F.; Boudinot, P.; Santos, Y.; MAGADAN, S.
Show abstract
Turbot (Scophthalmus maximus) is an important aquaculture species, but the genomic organization and expressed diversity of its antibody repertoire remain incompletely characterized. In this study, we annotated the immunoglobulin heavy chain (IGH) locus using the haplotype resolved fScoMax1.1 genome assembly, and we used this as a reference to profile the expressed turbot IgM, IgD and IgT repertoires in skin and spleen. The primary IGH locus was located on chromosome 19, spanned approximately 72 kb, and contained 25 IGHV genes, including 24 functional genes and one pseudogene, together with three IGHD, seven IGHJ and three IGHC genes corresponding to IgT, IgM and IgD. Comparison with the alternate fScoMax1.1 haplotype and a second turbot genome assembly showed conserved IGHD, IGHJ and IGHC content, whereas IGHV gene number differed among assemblies. High throughput 5RACE repertoire sequencing revealed isotype and tissue associated differences in expressed IGH diversity. IgM represented the dominant productive repertoire in both skin and spleen and showed the highest clonotypic diversity, particularly in spleen. IgD displayed an intermediate profile, whereas IgT was more enriched in skin and exhibited the strongest clonal restriction. IGHV subgroup usage was dominated by IGHV3 in IgM and IgD, whereas IgT showed a distinct profile characterized by preferential use of IGHV4, especially in skin. Gene level analysis further showed broad IGHV-IGHJ pairing in IgM and IgD, with preferential use IGHJ3 segment, while IgT sequences paired exclusively with IGHJT. Clonotype sharing between skin and spleen was isotype dependent, being strongest for IgT, intermediate for IgM, and negligible for IgD, suggesting that clonal expansion did not necessarily predict inter tissue trafficking. Together, these results provide a curated genomic and expressed repertoire framework for turbot IGH genes and reveal isotype specific organization of antibody diversity, with IgT displaying a particular repertoire pattern.
Pozo, G.; Rivas-Torres, G.; Velez-Darquea, E.; Barragan-Orbe, D.; Torres, M. d. L.
Show abstract
Scalesia gordilloi is a critically endangered species endemic to San Cristobal Island in the Galapagos archipelago and represents one of the most unique and vulnerable lineages within the adaptive radiation of the genus Scalesia. Despite its evolutionary distinctiveness and conservation importance, no genomic resources have been available for this species. Here, we present the first high-quality reference genome of S. gordilloi, generated using Oxford Nanopore long-read sequencing. Across three PromethION R10.4.1 flow cells, we obtained 80.5 Gb of long reads (~25X coverage), which enabled a highly contiguous 3.61 Gb assembly composed of only 549 contigs and an N50 of 106.6 Mb. BUSCO completeness reached 98.6%, with assembly metrics comparable to other high-quality Asteraceae genomes. Repeat annotation revealed that 76.2% of the genome is composed of interspersed elements, dominated by LTR retrotransposons. Structural annotation resulted in 47,913 high-confidence protein-coding genes, consistent with expectations for large, repetitive Asteraceae genomes. This genome provides a critical foundation for conservation genomics, enabling assessments of genetic diversity, inbreeding, and adaptive potential in the species. It further establishes a framework for comparative genomics across the Scalesia radiation and supports future efforts to protect and restore one of the most threatened plant lineages of the Galapagos Islands.
Rilwan, O.; Ibrahim, A.
Show abstract
Tomato (Solanum lycopersicum L.) is one of the most important vegetable crops in Nigeria, serving as a major source of income, nutrition, and raw material for food industries. However, its production is severely constrained by Fusarium wilt, a destructive soil-borne disease caused by Fusarium oxysporum f. sp. lycopersici. This study investigated the prevalence and severity of Fusarium wilt on tomato in Chikun Local Government Area (LGA) of Kaduna State, Nigeria. Field survey and laboratory analyses were conducted on forty-five tomato samples from three tomato farms Kujama, Kakau, and Rido. The samples were examined for disease incidence and severity. Data were analyzed using descriptive statistics and Chi-square tests. The overall disease incidence was with Rido recording the highest infection rate (80.0%), followed by Kujama (60.0%) and Kakau (40.0%). Among plant parts, the stem exhibited the highest infection frequency (80.0%), while leaves and fruits had 60.0% and 40.0% incidence respectively. Chi-square analysis indicated no significant difference (p > 0.05) in disease incidence among farms and plant parts, suggesting uniform pathogen distribution. The research recommends the adoption of integrated disease management strategies and improved farmer awareness to mitigate the impact of the disease and ensure sustainable tomato production.
Cocioba, S. S.; Huang, P.-C.; Mallon, J.; Chan, Z.; Geremew, A. W.; Bisson, A.; Kyriakakis, P.
Show abstract
Here we introduce OpenEvo, a fully open-source, low-cost turbidostat platform for automated continuous culture and directed evolution experiments. Existing tools are expensive, complex, or lack open-source hardware; OpenEvo addresses this gap. OpenEvo is a complete, fully automated evolution platform with detailed, illustrated construction instructions for beginners, open-source software and firmware, and a single device priced around $300. An optional PC-based version offers enhanced functionality, including remote access, programmable evolution cycles, programmable LED stimulation, and a data visualization tool. OpenEvo can cycle through three types of media for positive, negative, and neutral selection conditions, supporting a wide range of experimental designs. We validate the use of OpenEvo by evolving H. volcanii to grow from 15% to 12% salt over ~150 cycles, ~1,000 hours. Evolved cells grew 36% faster than wild-type at 12% salt. Whole-genome sequencing of adapted cells found SNPs and large deletions. We also demonstrate positive and negative selection using the OpenEvo LEDs to drive optogenetics via a Phytochrome B-based optogenetic tool, with light as the selection stimulus during over 4000 hours of growth. OpenEvo lowers the technical and cost barriers for continuous evolution experiments, serves as a teaching tool, and is designed to grow an open community of users who share modifications.
Miyamae, J. A.; Moore, T. Y.
Show abstract
Mammal tails have long been recognized for their diversity of morphological form and function, however, there remains a substantial gap between the motivation to understand and emulate the various performance functions of the tail and what is known about tail anatomy. In this study, we were motivated to discover the anatomical foundations of the fast, whipping motions of the tail of the lesser Egyptian jerboa (Jaculus jaculus), which may aid in the quick changes of direction as the animal escapes from predators using ricochetal bipedal hopping. We employed microCT scans, dissections, and museum data to describe the musculoskeletal anatomy of the jerboa in comparison with the laboratory mouse (Mus musculus) and rat (Rattus norvegicus). While many aspects of tail anatomy are conserved across these species, the jerboa does possess unique characteristics such as an extremely long tail arising from caudal vertebral elongation, development of extensive dorsal musculature differentiated into lateral and medial components to increase points of skeletal attachment, and a novel anatomical feature - the bi-lobed cranial transverse process - which serves as a supernumerary dorsal tendon attachment site and possible brace to protect the ventral tendons and intrinsic muscles for a section of caudal vertebrae which likely experiences high mechanical stress.
Xie, J.; Guo, Z.; Zhao, H.; Ni, H.
Show abstract
Abstract-Large language models (LLMs) [1], [2] have demon strated remarkable capabilities across general domains, yet their application in specialized medical contexts demands rigorous domain adaptation [3], [4]. We present Infoxmed2.0-27B, a medical foundation model built upon Qwen3.5-27B [5] through a comprehensive multi-stage post-training pipeline: (1) proprietary medical data synthesis from a MySQL database with MedicalCategoryTree organization, medical PhD team validation, Chinese RoBERTa [6] semantic deduplication, and API-assisted language refinement; (2) instruction supervised fine-tuning of Qwen3.5- 27B via LoRA [7] (r = 8, = 32) using MS-Swift [8], producing iterations Infoxmed2.0.0[->]2.0.2[->]2.0.4; (3) Direct Preference Optimization (DPO) [9] on 6,283 curated medical preference pairs [10] using DPO-RPO loss ({beta} = 0.3, RPO = 0.1) across eight progressive training iterations (v0-v7); and (4) parallel Group Relative Policy Optimization (GRPO) [11]-based medical reward model training on Qwen3.5 combining internal rule-based reward functions with external DeepSeek signals. Comprehensive evaluations under a uniform LLM-as-Judge [12] framework with GPT-5.4 demonstrate 77.0% accuracy (mean quality score +7.18) on MedMCQA [10] and +2.59 on HLE, with pipeline progression from +6.69 (base) to +7.06 (SFT) to +7.18 (final).
Qiu, X.; Wang, Y.; Wen, J.; Chen, Y.; Zhao, L.; Jian, J.; Yang, W.
Show abstract
The Wangs garden lizard, Calotes wangi, is a widely distributed agamid species in Southern China and Northern Vietnam and exhibits pronounced colour variation and rapid body colour change. Despite increasing interest in the genomic basis of colour variation, chromosome-level genomic resources remain limited in agamid lizards. Here, we generated a chromosome-level reference genome of C. wangi using PacBio HiFi sequencing and Hi-C scaffolding. The final genome assembly was approximately 1.66 Gb in size and comprised 6 macrochromosomes and 11 microchromosomes, with a contig N50 of 110.09 Mb and 98.9% complete BUSCO genes. A total of 20,442 protein-coding genes were annotated. Comparative genomic analyses identified 297 significantly expanded gene families, with enriched functions associated with steroid metabolism, chromatin regulation, and epigenetic processes. This high-quality genome assembly provides an important genomic resource for future studies of colour variation, phenotypic plasticity, and evolutionary diversification in agamid lizards.