Scientific Data
○ Springer Science and Business Media LLC
All preprints, ranked by how well they match Scientific Data's content profile, based on 209 papers previously published here. The average preprint has a 0.15% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Kleven, H.; Gillespie, T. H.; Zehl, L.; Dickscheid, T.; Bjaalie, J. G.; Martone, M. E.; Leergaard, T. B.
Show abstract
Brain atlases are important reference resources for accurate anatomical description of neuroscience data. Open access, three-dimensional atlases serve as spatial frameworks for integrating experimental data and defining regions-of-interest in analytic workflows. However, naming conventions, parcellation criteria, area definitions, and underlying mapping methodologies differ considerably between atlases and across atlas versions. This lack of standardization impedes use of atlases in analytic tools and registration of data to different atlases. To establish a machine-readable standard for representing brain atlases, we identified four fundamental atlas elements, defined their relations, and created an ontology model. Here we present our Atlas Ontology Model (AtOM) and exemplify its use by applying it to mouse, rat, and human brain atlases. We propose minimum requirements for FAIR atlases and discuss how AtOM may facilitate atlas interoperability and data integration. AtOM provides a standardized framework for communication and use of brain atlases to create, use, and refer to specific atlas elements and versions. We argue that AtOM will accelerate analysis, sharing, and reuse of neuroscience data.
Nwosu, I. O.; Tabler, D. D.; Chipman, G.; Piccolo, S. R.
Show abstract
Transcriptomic data from breast-cancer patients are widely available in public repositories. However, before a researcher can perform statistical inferences or make biological interpretations from such data, they must find relevant datasets, download the data, and perform quality checks. In many cases, it is also useful to normalize and standardize the data for consistency and to use updated genome annotations. Additionally, researchers need to parse and interpret metadata: clinical and demographic characteristics of patients. Each of these steps requires computational and/or biomedical expertise, thus imposing a barrier to reuse for many researchers. We have identified and curated 102 publicly available, breast-cancer datasets representing 17,151 patients. We created a reproducible, computational pipeline to download the data, perform quality checks, renormalize the raw gene-expression measurements (when available), assign gene identifiers from multiple databases, and annotate the metadata against the National Cancer Institute Thesaurus, thus making it easier to infer semantic meaning and compare insights across datasets. We have made the curated data and pipeline freely available for other researchers to use. Having these resources in one place promises to accelerate breast-cancer research, enabling researchers to address diverse types of questions, using data from a variety of patient populations and study contexts.
Hadar, P. N.; Li, J.; Coughlin, B.; Munoz, W.; Hsueh, B.; Williams, Z. M.; Yee, S.; Rapalino, O.; Brandman, D.; Stavisky, S. D.; Henderson, J. M.; Willett, F. R.; Rubin, D. B.; Hochberg, L. R.; Cash, S. S.; Choi, E. Y.; Paulk, A. C.
Show abstract
As implantable brain-computer interfaces (iBCIs) for communication and movement transition from cutting-edge research to clinical practice, a standardized approach will be required to reliably plan neurosurgeries involving complex microelectrode arrays and other neural sensors. Here, through our BrainGate study experiences, we present a replicable methodology, using open-source tools, to create interactive, personalized, 3-dimensional, virtual and physical, functional mapping models to guide iBCI surgical planning and provide intra-operative imaging displays.
Mou, X.; He, C.; Tan, L.; Yu, J.; Liang, H.; Zhang, J.; Tian, Y.; Yang, Y.; Xu, T.; Wang, Q.; Cao, M.; Chen, Z.; Hu, C.; Wang, X.; Liu, Q.; Wu, H.
Show abstract
An Electroencephalography (EEG) dataset utilizing rich text stimuli can advance the understanding of how the brain encodes semantic information and contribute to semantic decoding in brain-computer interface (BCI). Addressing the scarcity of EEG datasets featuring Chinese linguistic stimuli, we present the ChineseEEG dataset, a high-density EEG dataset complemented by simultaneous eye-tracking recordings. This dataset was compiled while 10 participants silently read approximately 11 hours of Chinese text from two well-known novels. This dataset provides long-duration EEG recordings, along with pre-processed EEG sensor-level data and semantic embeddings of reading materials extracted by a pre-trained natural language processing (NLP) model. As a pilot EEG dataset derived from natural Chinese linguistic stimuli, ChineseEEG can significantly support research across neuroscience, NLP, and linguistics. It establishes a benchmark dataset for Chinese semantic decoding, aids in the development of BCIs, and facilitates the exploration of alignment between large language models and human cognitive processes. It can also aid research into the brains mechanisms of language processing within the context of the Chinese natural language.
Rampinini, A. C.; Balboni, I.; Kepinska, O.; Berthele, R.; Golestani, N.
Show abstract
This paper introduces the "NEBULA101 - Neuro-behavioural Understanding of Language Aptitude" dataset, which comprises behavioural and brain imaging data from 101 healthy adults to examine individual differences in language and cognition. Human language, a multifaceted behaviour, varies significantly among individuals, at different processing levels. Recent advances in cognitive science have embraced an integrated approach, combining behavioural and brain studies to explore these differences comprehensively. The NEBULA101 dataset offers brain structural, diffusion-weighted, task-based and resting-state MRI data, alongside extensive linguistic and non-linguistic behavioural measures to explore the complex interaction of language and cognition in a highly multilingual sample. By sharing this multimodal dataset, we hope to promote research on the neuroscience of language, cognition and multilingualism, enabling the field to deepen its understanding of the multivariate panorama of individual differences and ultimately contributing to open science.
Wenk, E. H.; Sauquet, H.; Gallagher, R. V.; Brownlee, R.; Boettiger, C.; Coleman, D.; Yang, S.; Auld, T.; Barrett, R. L.; Brodribb, T.; Choat, B.; Dun, L.; Ellsworth, D.; Gosper, C.; Guja, L.; Jordan, G. J.; Breton, T.; Leigh, A.; Irving, P.; Medlyn, B.; Nolan, R.; Ooi, M.; Sommerville, K. D.; Vesk, P.; White, M.; Wright, I. J.; Falster, D. S.
Show abstract
Traits with intuitive names, a clear scope and explicit description are essential for all trait databases. Reanalysis of data from a single database, or analyses that integrate data across multiple databases, can only occur if researchers are confident the trait concepts are consistent within and across sources. The lack of a unified, comprehensive resource for plant trait definitions has previously limited the utility of trait databases. Here we describe the AusTraits Plant Dictionary (APD), which extends the trait definitions included in the new trait database AusTraits. The development process of the APD included three steps: review and formalisation of the scope of each trait and the accompanying trait description; addition of trait meta-data; and publication in both human and machine-readable forms. Trait definitions include keywords, references and links to related trait concepts in other databases, and the traits are grouped into a hierarchy for easy searching. As well as improving the usability of AusTraits, the Dictionary will foster the integration of trait data across global and regional plant trait databases.
Maschke, C.; Hadar, P. N.; Zhang, Y.; Li, J.; Ganjoo, G.; Hoopes, A.; Guazzo, A.; Gupta, A.; Ghanta, M.; Nearing, B.; Silvers, C. T.; Gunapati, B.; Thomas, R.; Kim, J. A.; Mukerji, S. S.; Dalca, A.; Zafar, S.; Lam, A.; Mignot, E.; Westover, M. B.
Show abstract
1The Brain Imaging and Neurophysiology Database (BIND) represents one of the largest multi-institutional, multimodal, clinical neuroimaging repositories, comprising 1.8 million brain scans from 38,945 patients, linked to neurophysiological recordings. This comprehensive dataset addresses critical limitations in neuroimaging research by providing unprecedented scale and diversity across pathologies and health. BIND integrates de-identified data from Massachusetts General Hospital, Brigham and Womens Hospital, and Stanford University, including 1,723,699 MRI scans (1.5 Tesla, 3 Tesla, and 7 Tesla), 54,137 CT scans, 5,093 PET scans, and 526 SPECT scans, converted to standardized NIfTI format following BIDS organization. The database spans the full age spectrum (newborn to 106 years) and encompasses diverse neurological conditions alongside healthy patients. We deployed Bio-Medical Large Language Models to extract structured clinical metadata from 84,960 brain-related reports, categorizing findings into standardized pathology classifications. All imaging data are linked to previously published EEG and polysomnography recordings from the Harvard Electroencephalography Database, enabling unprecedented multimodal analyses. BIND is freely accessible for academic research through the Brain Data Science Platform (https://bdsp.io/). This resource facilitates large-scale neuroimaging studies, machine learning applications, and multimodal brain research to accelerate discoveries in clinical neuroscience.
Lan, Z.; Zhang, T.; Zhang, Y.; Sun, X.; Liu, C.; Liu, P.; Tang, M.; Fu, M.; Hagood, J. S.; Pickles, R. J.; Zou, F.; Zheng, X.
Show abstract
Our understanding of the human immune systems response to viral respiratory tract infections (VRTIs) and vaccines, including the molecular mechanisms and correlates of protection, remains incomplete. Extensive transcriptomic data from inoculation and vaccination studies have been deposited in publicly available databases. However, these studies are often separate and difficult to locate. We provide a curated compendium of public gene expression data repositories for researchers to reanalyze transcriptomes from human whole blood, peripheral blood mononuclear cells (PBMCs), and nasal swab samples. This enables the study of transcriptional responses to viral inoculation or vaccination. This collection includes 18 datasets from inoculation studies and 37 datasets from vaccination studies, sourced from the NCBI Gene Expression Omnibus (GEO), ImmPort Shared Data, and ArrayExpress.
Ramesh, V.; Singh, S.; Pop, P.; Choksi, P.; Singh, P.; Khanwilkar, S.; Teotia, S.; Burli, P.; Devarajan, K.; A, A.; A Nakhwa, A.; Abdus Shakur, M.; Baishya, R.; Bhagwat, N.; Biniwale, S.; Bora, C.; C S, S.; Chakraborty, N.; D'Souza, S.; D'Souza, E.; Vaishnav, R. D.; Deshpande, K.; Dhanda, A.; G, A.; Ghosh, A.; Goswami, R.; K N, A.; K P, N.; K Rajaraman, B.; K V, G.; Kannan, V.; Karthick, V.; Kotian, M.; Kumar, H.; Kurian, P.; Madhavan, M.; Meena, K.; Mohammad Maslehuddin, A.; Mourya, P.; Mudke, M.; R J, P.; R S Jha, R.; Ramesh, K.; Sailas, S. S.; Sangwan, T.; Mahesh, S.; Satish, R.; Shankar, A
Show abstract
Global rates of biodiversity loss warrant conservation action and monitoring at large geographic scales. Conservation technologies such as acoustic monitoring in conjunction with deep learning now enable us to monitor wildlife simultaneously across space and time. However, for a significant proportion of biodiversity in tropical regions, we cannot yet rely on automated recognition approaches because we lack acoustic templates to robustly train deep learning algorithms. In this paper, we relied on a novel participatory approach, enlisting researchers, conservation practitioners, and nature enthusiasts to create a unique crowd-sourced, open-access dataset of acoustic annotations across taxonomic groups for biodiversity in India. Our dataset comprises 3311 minutes of strongly labelled data (bounding boxes or annotations for a species vocalization) and 2504 minutes of weakly labelled data (indicating the presence of a species within an audio file but lacking bounding boxes) for 518 species across India, spanning 25 of 36 states and union territories. We present metadata and code for data processing and highlight the strengths of a participatory approach to biodiversity monitoring.
Manfrini, E.; Sauvion, N.; Maquart, P.-O.; Legal, L.; Blight, O.; Duquesne, E.; Hanot, C.; Bang, A.; Geslin, B.; Goebel, F.-R.; Fournier, D.; Berggren, A.; Javal, M.; Angulo, E.; Pincebourde, S.; Zakardjian, M.; Renault, D.; Le Lann, C.; Derocles, S.; Vayssieres, J.-F.; Leroy, B.; Courchamp, F.
Show abstract
Insect research remains hindered by limited data availability and fragmented knowledge compared to other, better-documented taxonomic groups. Increasingly, both the macroecological and the insect research communities highlight the need to integrate large-scale ecological trait datasets for insects. We present AnthropInsect, the largest database on insect traits to date, which uniquely includes variables describing human-insect associations. AnthropInsect describes species through 35 variables grouped into five categories: (i) taxonomic descriptors; (ii) ecological descriptors (native bioregions and habitat); (iii) human-insect associations (edibility and invasive status); (iii) functional traits (behavior, morphology, life history and feeding); (iv) and macroecological descriptors of native-range geography and climate. AnthropInsect currently includes 5,870 species across six major orders: Coleoptera, Lepidoptera, Hemiptera, Hymenoptera, Orthoptera and Blattodea. Data extracted from peer-reviewed and grey literature and from existing databases were standardized and curated with expert knowledge to ensure accuracy. By providing traits data with information on insect- human interactions, this rigorously curated resource supports global research in entomology, ecology, conservation, and global change.
McCoy, L. G.; Smith, J.; Anchuri, K.; Berry, I.; Pineda, J.; Harish, V.; Lam, A. T.; Yi, S. E.; Hu, S.; Canadian Open Data Working Group: Non-Pharmaceutical Interventions, ; Fine, B.
Show abstract
Non-pharmaceutical interventions (NPIs) have been the primary tool used by governments and organizations to mitigate the spread of the ongoing pandemic of COVID-19. Natural experiments are currently being conducted on the impact of these interventions, but most of these occur at the subnational level - data not available in early global datasets. We describe the rapid development of the first comprehensive, labelled dataset of 1640 NPIs implemented at federal, provincial/territorial and municipal levels in Canada to guide COVID-19 research. For each intervention, we provide: a) information on timing to aid in longitudinal evaluation, b) location to allow for robust spatial analyses, and c) classification based on intervention type and target population, including classification aligned with a previously developed measure of government response stringency. This initial dataset release (v1.0) spans January 1st, and March 31st, 2020; bi-weekly data updates to continue for the duration of the pandemic. This novel dataset enables robust, inter-jurisdictional comparisons of pandemic response, can serve as a model for other jurisdictions and can be linked with other information about case counts, transmission dynamics, health care utilization, mobility data and economic indicators to derive important insights regarding NPI impact.
Gomes, P. W. P.; Mannochio-Russo, H.; Schmid, R.; Zuffa, S.; Damiani, T.; Quiros-Guerrero, L.-M.; Caraballo-Rodriguez, A. M.; Zhao, H. N.; Yang, H.; Xing, S.; Charron-Lamoureux, V.; Chigumba, D. N.; Sedio, B. E.; Myers, J. A.; Allard, P.-M.; Harwood, T. V.; Tamayo-Castillo, G.; Kang, K. B.; Defossez, E.; Koolen, H. H. F.; da Silva, M. N.; e Silva, C. Y. Y.; Rasmann, S.; Walker, T. W. N.; Glauser, G.; Chaves-Fallas, J. M.; David, B.; Kim, H.; Lee, K. H.; Kim, M. J.; Choi, W. J.; Keum, Y.-S.; de Lima, E. J. S. P.; de Medeiros, L. S.; Bataglion, G. A.; Costa, E. V.; da Silva, F. M. A.; Carvalho,
Show abstract
Understanding the distribution of hundreds of thousands of plant metabolites across the plant kingdom presents a challenge. To address this, we curated publicly available LC-MS/MS data from 19,075 plant extracts and developed the plantMASST reference database encompassing 246 botanical families, 1,469 genera, and 2,793 species. This taxonomically focused database facilitates the exploration of plant-derived molecules using tandem mass spectrometry (MS/MS) spectra. This tool will aid in drug discovery, biosynthesis, (chemo)taxonomy, and the evolutionary ecology of herbivore interactions.
Wilhelm, L.; Wang, Y.; Xu, S.
Show abstract
The Colorado potato beetle (CPB) is a major pest of potato crops that has evolved resistance to more than 50 pesticides. For decades, CPB has been a model species for research on insecticide resistance, insect physiology, diapause, reproduction and evolution. Yet, the research progress in CPB is constrained by the lack of comprehensive genomic and transcriptomic information. Here, building on the recently established chromosome-level genome assembly, we built a gene expression atlas of the CPB using the transcriptomes of 61 samples representing major organs and developmental stages. By using both short and long reads, we improved the genome annotation and identified 6,658 more genes that were missed in previous annotations. We then established a web portal allowing the search and visualization of the gene expression for the research community. The CPB atlas provides useful tools and comprehensive gene expression data, which will accelerate future research in both pest control and insect biology fields.
Novelli, L.; Stoliker, D.; Barta, T.; Greaves, M. D.; Chopra, S.; Jackson, J.; Kwee, J.; Williams, M.; Razi, A.
Show abstract
PsiConnect is a large-scale neuroimaging study designed to investigate the neural and subjective effects of psilocybin using multimodal neuroimaging. It combines functional, structural, and diffusion-weighted MRI with EEG to examine brain activity in 62 participants before and after a 19 mg dose of psilocybin. The design includes resting-state scans and three naturalistic conditions: guided meditation, music listening, and movie watching. Half of the cohort underwent an 8-week meditation training program, enabling the exploration of interactions among meditation, psilocybin, and brain function. The fMRI data was obtained through multi-echo fMRI, which enhances the signal-to-noise ratio and reduces susceptibility artifacts, thereby improving the reliability of the analyses. A comprehensive battery of behavioural and self-report measures captured both acute and longitudinal cognitive and subjective effects, with follow-ups extending to one year post-administration. The large sample size, multimodal neuroimaging, diversity of contexts, and longitudinal behavioural follow-ups enable the study of psilocybin-induced changes in brain and behaviour with an unprecedented level of detail and reliability. Furthermore, the data is curated according to open science principles to ensure accessibility and interoperability with established neuroimaging processing pipelines. These factors make PsiConnect a valuable and highly reusable resource for researchers in cognitive and computational neuroscience.
Clarke, N.; Wang, H.-T.; Lussier, D.; Bore, A.; Tetrel, L.; Beaudoin, C.; Das, S.; Pilon, R.; Evans, A. C.; Chertkow, H.; Dixon, R. A.; Badhwar, A.; Duchesne, S.; Bellec, L.
Show abstract
Resting-state functional connectivity (RSFC) holds promise for the detection and characterisation of dementia. The Comprehensive Assessment of Neurodegeneration and Dementia (COMPASS-ND) Study, by the Canadian Consortium on Neurodegeneration in Aging (CCNA), provides a unique resource to study deeply phenotyped neurodegenerative conditions. We present RSFC derivatives for 784 participants (data release 7 of the cohort) who were either cognitively unimpaired or diagnosed primarily with Alzheimers disease (AD), mixed dementia (AD with a vascular component), mild cognitive impairment (MCI), vascular MCI, frontotemporal dementia, Parkinsons disease with or without MCI or dementia, Lewy body disease or subjective cognitive impairment. Functional MRI scans were preprocessed using fMRIPrep, and time-series and whole-brain connectomes generated using three atlases at multiple resolutions, denoised using seven different techniques. High-motion artifacts were managed using a liberal quality control threshold appropriate for an older clinical population, resulting in data from 680 participants. These derivatives are made available to the research community to accelerate research on RSFC biomarkers of neurodegenerative disease, reducing duplication of effort, saving computational resources, and improving standardisation across studies.
Li, J.; Bhattasali, S.; Zhang, S.; Franzluebbers, B.; Luh, W.-M.; Spreng, R. N.; Brennan, J. R.; Yang, Y.; Pallier, C.; Hale, J. T.
Show abstract
Neuroimaging using more ecologically valid stimuli such as audiobooks has advanced our understanding of natural language comprehension in the brain. However, prior naturalistic stimuli have typically been restricted to a single language, which limited generalizability beyond small typological domains. Here we present the Le Petit Prince fMRI Corpus (LPPC-fMRI), a multilingual resource for research in the cognitive neuroscience of speech and language during naturalistic listening (Open-Neuro: ds003643). 49 English speakers, 35 Chinese speakers and 28 French speakers listened to the same audiobook The Little Prince in their native language while multi-echo functional magnetic resonance imaging was acquired. We also provide time-aligned speech annotation and word-by-word predictors obtained using natural language processing tools. The resulting timeseries data are shown to be of high quality with good temporal signal-to-noise ratio and high inter-subject correlation. Data-driven functional analyses provide further evidence of data quality. This annotated, multilingual fMRI dataset facilitates future re-analysis that addresses cross-linguistic commonalities and differences in the neural substrate of language processing on multiple perceptual and linguistic levels.
Bandrowski, A.; Grethe, J. S.; Pilko, A.; Gillespie, T. H.; Pine, G.; Patel, B.; Surles-Zeiglera, M.; Martone, M. E.
Show abstract
The NIH Common Funds Stimulating Peripheral Activity to Relieve Conditions (SPARC) initiative is a large-scale program that seeks to accelerate the development of therapeutic devices that modulate electrical activity in nerves to improve organ function. Integral to the SPARC program are the rich anatomical and functional datasets produced by investigators across the SPARC consortium that provide key details about organ-specific circuitry, including structural and functional connectivity, mapping of cell types and molecular profiling. These datasets are provided to the research community through an open data platform, the SPARC Portal. To ensure SPARC datasets are Findable, Accessible, Interoperable and Reusable (FAIR), they are all submitted to the SPARC portal following a standard scheme established by the SPARC Curation Team, called the SPARC Data Structure (SDS). Inspired by the Brain Imaging Data Structure (BIDS), the SDS has been designed to capture the large variety of data generated by SPARC investigators who are coming from all fields of biomedical research. Here we present the rationale and design of the SDS, including a description of the SPARC curation process and the automated tools for complying with the SDS, including the SDS validator and Software to Organize Data Automatically (SODA) for SPARC. The objective is to provide detailed guidelines for anyone desiring to comply with the SDS. Since the SDS are suitable for any type of biomedical research data, it can be adopted by any group desiring to follow the FAIR data principles for managing their data, even outside of the SPARC consortium. Finally, this manuscript provides a foundational framework that can be used by any organization desiring to either adapt the SDS to suit the specific needs of their data or simply desiring to design their own FAIR data sharing scheme from scratch.
McGeoch, M. A.; Pagad, S.; Bisset, S.; Genovesi, P.; Groom, Q.; Hirsch, T.; Jetz, W.; Ranipeta, A.; Schigel, D.; Sica, Y. V.
Show abstract
The Country Compendium of the Global Register of Introduced and Invasive Species (GRIIS) is a collation of data across 196 individual country checklists of alienspecies, along with a designation of those species with evidence of impact at a country level. The Compendium provides a baseline for monitoring the distribution and invasion status of all major taxonomic groups, and can be used for the purpose of global analyses of Introduced (alien, non-native, exotic) and Invasive species (invasive alien species), including regional, single and multi-species taxon assessments and comparisons. It enables exploration of gaps and inferred absences of species for countries, and provides a means for short to medium term refinement of GRIIS Checklists. The Country Compendium is, for example, instrumental, along with data on first records of introduction, for assessing and reporting on invasive alien species targets, including for the Convention on Biological Diversity and Sustainable Development Goals. The GRIIS Country Compendium provides a baseline and mechanism for tracking the spread of introduced and invasive alien species across countries globally. O_TBL View this table: org.highwire.dtl.DTLVardef@eb30fdorg.highwire.dtl.DTLVardef@dd338aorg.highwire.dtl.DTLVardef@62b8daorg.highwire.dtl.DTLVardef@1561763org.highwire.dtl.DTLVardef@119988e_HPS_FORMAT_FIGEXP M_TBL C_TBL
Pedro A. Valdes-Sosa; Jorge F. Bosch-Bayard; Lidice Galan Garcia; Maria L Bringas Vega; Eduardo Aubert Vazquez; Samir Das; Trinidad Virues Alba; Cecile Madjar; Zia Mohades; Leigh C. MacIntyre; Chrystine Rogers; Shawn Brown; Lourdes Valdes Urrutia; Iris Rodriguez Gil; Alan C. Evans; Mitchell J. Valdes Sosa
Show abstract
The Cuban Human Brain Mapping Project (CHBMP) repository is an open multimodal neuroimaging and cognitive dataset from 282 healthy participants, age range 18 to 68 years (mean 31.9 SD 9.3 years). This dataset was acquired from 2004 to 2008 as a subset of a larger stratified random sample of 2,019 participants from La Lisa municipality in La Habana, Cuba. The exclusion included presence of disease or brain dysfunctions. The information made available for all participants comprises: high-density (64-120 channels) resting state electroencephalograms (EEG), magnetic resonance images (MRI), psychological tests (MMSE, Wechsler Adult Intelligence Scale WAIS III, computerized reaction time tests using a go no-go paradigm), as well as general information (age, gender, education, ethnicity, handedness and weight). The EEG data contains recordings with at least 30 minutes duration including the following conditions: eyes closed, eyes open, hyperventilation and subsequent recovery. The MRI consisted in anatomical T1 and T2 as well as diffusion weighted (DWI) images acquired on a 1.5 Tesla system. The data is available for registered users on the LORIS database which is part of the MNI neuroinformatics ecosystem.Competing Interest StatementThe authors have declared no competing interest.
Yamaguchi, K.; Hara, Y.; Kaori, T.; Nishimura, O.; Smith, J.; Kadota, M.; Kuraku, S.
Show abstract
The group of hagfishes (Myxiniformes) arose from agnathan (jawless vertebrate) lineages and is one of the only two extant cyclostome taxa, together with lampreys (Petromyzontiformes). Even though whole genome sequencing has been achieved for diverse vertebrate taxa, genome-wide sequence information has been highly limited for cyclostomes. Here we sequenced the genome of the inshore hagfish Eptatretus burgeri using DNA extracted from the testis, with a short-read sequencing platform, aiming at reconstructing a high-coverage coding gene catalogue. The obtained genome assembly, scaffolded with mate-pair reads and paired RNA-seq reads, exhibited an N50 scaffold length of 293 Kbp, which allowed the genome-wide prediction of coding genes. This computation resulted in the gene models whose completeness was estimated at the complete coverage of more than 83 % and the partial coverage of more than 93 % by referring to evolutionarily conserved single-copy orthologs. The high contiguity of the assembly and completeness of resulting gene models promises a high utility in various comparative analyses including phylogenomics and phylome exploration.