Database
◐ Oxford University Press (OUP)
All preprints, ranked by how well they match Database's content profile, based on 61 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Yu, W.; Gwinn, M.; Khoury, M. J.
Show abstract
SummaryWe developed a new online database that contains the most updated published scientific literature, online news and reports, CDC and National Institutes of Health (NIH) resources. The tool captures emerging discoveries and applications of genomics, molecular, and other precision medicine and precision public health tools in the investigation and control of coronavirus diseases, including COVID-19, MERS-CoV, and SARS. AvailabilityCoronavirus Disease Portal (CDP) can be freely accessed via https://phgkb.cdc.gov/PHGKB/coVInfoStartPage.action. Contactwyu@cdc.gov
Soundarajan, S.; Kuruppu, S.; Singh, A.; Kim, J.; Achalla, M.
Show abstract
The NIH SPARC program seeks to accelerate the development of therapeutic devices that modulate electrical activity in nerves to improve organ function. SPARC-funded researchers are generating rich datasets from neuromodulation research that are curated and shared according to FAIR (Findable, Accessible, Interoperable, and Reusable) guidelines and are accessible to the public on the SPARC data portal. Keeping track of the utilization of these datasets within the larger research community is a feature that will benefit data generating researchers in showcasing the impact of their SPARC outcomes. This will also allow the SPARC program to display the impact of the FAIR data curation and sharing practices that have been implemented. This manuscript provides the methods and outcomes of SPARClink, our web tool for visualizing the impact of SPARC, which won the 2nd prize at the 2021 SPARC FAIR Codeathon. With SPARClink, we built a system that automatically and continuously finds new published SPARC scientific outputs (datasets, publications, protocols) and the external resources referring to them. SPARC datasets and protocols are queried using publicly accessible REST APIs (provided by Pennsieve and Protocols.io) and stored in a publicly accessible database. Citation information for these resources is retrieved using the NIH reporter API and NCBI Entrez system. A novel knowledge-graph-based structure was created to visualize the results of these queries and showcase the impact that the FAIR data principles can have on the research landscape when they are adopted by a consortium.
Cano, M. A.; Tsueng, G.; Zhou, X.; Hughes, L. D.; Mullen, J. L.; Xin, J.; Su, A. I.; Wu, C.
Show abstract
BackgroundBiomedical researchers are strongly encouraged to make their research outputs more Findable, Accessible, Interoperable, and Reusable (FAIR). While many biomedical research outputs are more readily accessible through open data efforts, finding relevant outputs remains a significant challenge. Schema.org is a metadata vocabulary standardization project that enables web content creators to make their content more FAIR. Leveraging schema.org could benefit biomedical research resource providers, but it can be challenging to apply schema.org standards to biomedical research outputs. We created an online browser-based tool that empowers researchers and repository developers to utilize schema.org or other biomedical schema projects. ResultsOur browser-based tool includes features which can help address many of the barriers towards schema.org-compliance such as: The ability to easily browse for relevant schema.org classes, the ability to extend and customize a class to be more suitable for biomedical research outputs, the ability to create data validation to ensure adherence of a research output to a customized class, and the ability to register a custom class to our schema registry enabling others to search and re-use it. We demonstrate the use of our tool with the creation of the Outbreak.info schema--a large multi-class schema for harmonizing various COVID-19 related resources. ConclusionsWe have created a browser-based tool to empower biomedical research resource providers to leverage schema.org classes to make their research outputs more FAIR.
Asaduzzaman, S.; Bansal, B.; Combs, P.; Zhang, J.; Rehana, H.; McGregor, B.; He, Y.; Hur, J.
Show abstract
BackgroundThe expansion of biomedical literature demands systematic ontology-guided discovery of gene interactions, vaccine mechanisms, drug associations, and adverse events. Existing platforms such as STRING, DisGeNET, and PubTator fall short of providing a unified, freely accessible system that integrates ontology-based semantic interaction classification, vaccine-focused heterogeneous network construction, and Artificial Intelligence-assisted evidence retrieval. ResultsIgnet 2.0 and Vignet are freely accessible dual-platform systems that combine PubMed literature mining, BioBERT-based interaction scoring for millions of gene-gene co-occurrence pairs and integrate three biomedical ontologies and one curated drug resource, Interaction Network Ontology (INO), Vaccine Ontology (VO), Human Disease Ontology (HDO), and DrugBank. Ignet 2.0 supports gene interaction discovery, gene set enrichment retrieval of BioBERT-scored GenePair evidence, and AI-assisted summarization through BioSummarAI. Vignet extends these features with VO-guided Vaccine Exploration, VacPair interaction scoring, and the creation of vaccine, gene, drug, and disease networks in VacNet. A public Representational State Transfer Application Programming Interface (REST API) and Model Context Protocol (MCP) endpoint enable real-time integration, fostering trust in biomedical knowledge discovery. ConclusionIgnet 2.0 and Vignet are scalable, ontology-guided biomedical knowledge platforms that facilitate evidence-based gene interaction analysis, vaccine-focused semantic exploration, and AI-assisted knowledge discovery. Their real-time PubMed data integration ensures up-to-date insights; however, users should consider validation processes and potential lags in incorporating the latest experimental data, which may affect the reliability of immediate data. AvailabilityIgnet 2.0: https://ignet.org/ignet; Vignet: https://ignet.org/vignet/
Du, Y.; Ouyang, W.; Kitzmiller, J.; Guo, M.; Zhao, S.; Whitsett, J. A.; Xu, Y.
Show abstract
Recent advances in single-cell omics and high-resolution imaging have provided unanticipated data resources for the elucidation of genes underlying the complex biological processes critical for organ formation and function. However, processing and integrating large amounts of single-cell omics and imaging data presents a major challenge for most researchers. There is a critical need for ready-to-use computational tools for data/knowledge integration and visualization. Here we present "Lung-at-a-glance", an easy-to-use web toolset for visualizing and interoperating complex omics and imaging data, providing an interactive web interface to bridge lung anatomic ontology classifications to lung histology and immunofluorescence confocal images, and cell-type-specific gene expression. "Lung-at-a-glance" contains three interactive components: 1) "Region at a glance", 2) "Cell at a glance" and 3) "Gene at a glance". "Lung-at-a-glance" and other newly developed web tools for lung-related data query, integration and visualization are publicly available on LGEA web portal v3 https://research.cchmc.org/pbge/lunggens/mainportal.html.
Zhang, D.; Sun, H.-X.; Zhou, Z.; Jiang, X.; Chen, D.; Zhou, S.; Huang, J.; Qu, S.; Gu, Y.; Zhang, X.; Jin, X.; Gao, Y.; Shen, Y.; Chen, F.
Show abstract
Birth defect, not only poses a major challenge for infant health but also attracts the attention of countless people in the world. Chromosome abnormality directly results in diverse birth defects which are generally deleterious and even lethal. Therefore, gaining molecular regulatory insights into these diseases is important and necessary for effective prenatal screening. Recently, with the advance of next-generation sequencing (NGS) techniques, a myriad of treatises and data associated with these diseases are now constantly produced from different laboratories across the world. To meet the increasing requirements for birth-related data resources, we developed a birth defect multi-omics database (BDdb), freely accessible at http://t21omics.cngb.org and consisting of multi-omics data, circulating free DNA (cfDNA) data, as well as diseases biomarkers. Omics data sets from 138 GSE samples, 5271 GSM samples and 328 entries, and more than 2000 biomarkers of 22 birth-defect diseases in 5 different species were integrated into BDdb, which provides a user-friendly interface for searching, browsing and downloading selected data. Additionally, we re-analyzed and normalized the raw data so that users can also customize the analysis using the data generated from different sources or different High-Throughput Sequencing (HTS) methods. To our knowledge, BDdb is the first comprehensive database associated with birth-defect-related diseases. which would benefit the diagnosis and prevention of birth defects.
Pinero, J.; Corvi, J.; Rykova, N.; Guillem, A.; Martinez, A.; Zufiaur, J. T.; Rivetta, I.; Holmes, S.; Del Boccio, M.; Shapoval, D.; Szneiderowicz, M. D. S.; Slager, F.; Sanz, F.; Furlong, L. I.
Show abstract
Precision medicine and therapeutic development rely on a comprehensive understanding of genotype-phenotype relationships, yet this information remains fragmented across diverse sources. DISGENET, established over 15 years ago, addresses this challenge by systematically integrating gene-disease, variant-disease, and disease-disease associations from authoritative databases and the literature. This major upgrade expands coverage with chemical and pharmacological annotations and integrates biobank and clinical data. An advanced natural language processing (NLP) pipeline captures emerging evidence with full provenance and contextual details, key to streamlining data-driven insights. DISGENET supports diverse users through multiple tools, including an intuitive web interface, a REST API, an R package, a Cytoscape app, and an AI assistant for natural language queries. Quarterly updates ensure data currency, while a sustainable freemium model provides free academic access and supports ongoing development. DISGENET aims to accelerate data-driven discoveries and advance precision medicine and drug development. The platform is accessible at https://www.disgenet.com. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=80 SRC="FIGDIR/small/697749v1_ufig1.gif" ALT="Figure 1"> View larger version (33K): org.highwire.dtl.DTLVardef@1cf5b9aorg.highwire.dtl.DTLVardef@87004aorg.highwire.dtl.DTLVardef@1243b2borg.highwire.dtl.DTLVardef@1a8adab_HPS_FORMAT_FIGEXP M_FIG C_FIG HighlightsO_LIDISGENET accelerates precision medicine via comprehensive genotype-phenotype data C_LIO_LIAccess to gene, variant, disease, and drug data in a single platform C_LIO_LIUp-to-date evidence with provenance and full context C_LIO_LISuite of tools to serve diverse research and clinical communities C_LIO_LISustainable freemium model ensures ongoing platform innovation. C_LI
Lain, A. D.; Go, S.; Mahmud, A.; Rajendra, S.; Cano San Jose, A.; Loupasaki, K.; Theodoridis, G.; Bizkarguenaga Uribiarte, M.; Gu, Y.; Deda, O.; Conde, R. D. A.; Embade, N.; de Diego Rodriguez, A.; Burguera, N.; Rossiou, D.; Gil Redondo, R.; Gallou, D.; Tueros, I.; Velmurugan, R.; Gkanali, V.; Caro Burgos, M.; Pousinis, P.; Alektoridis, G.; Arranz, S.; Nikolopoulos, N.; Yan, X.; Fernandez Carrion, R.; Rowlands, T.; Choi, D.; Rei, M.; Cave-Ayland, C.; D Alessandro, A.; The CoDiet consortium, ; Beck, T.; Posma, J. M.
Show abstract
We present here four biomedical, multi-entity corpora that can be used as benchmarks for named-entity recognition (NER), targeted to literature on metabolic syndrome. The CoDiet-Gold corpus (348,413 annotations) contains 500 re-distributable full-text publications, of which each document was independently annotated by two human experts, with disagreements fully adjudicated by a third expert. The CoDiet-Electrum corpus (2,349,499 annotations) contains 3,688 publications that were annotated using the entities from CoDiet-gold. Finally, for the same 3,688 documents, two fully machine-annotated corpora CoDiet-Bronze (2,399,647 annotations) and CoDiet-Silver (1,868,422 annotations), were created by utilising existing NER algorithms to annotate these. These corpora contain categories (organisms, disease, genes, proteins, metabolites) that add depth to existing corpora, as well as new categories that do not have other corpora (food, dietary methods, sample types, computational methods, study methodology, population characteristics, data types, and microbiome).
Mohseni Ahooyi, T.; Stear, B. J.; Taylor, D. M.
Show abstract
The Homo sapiens Chromosomal Location Ontology for GRCh38 (HSCLO38) represents a knowledge-graph-ready framework for connecting genomic features at multiple resolutions. We present the methodology behind the development of HSCLO38 and its integration with current genomic standards for application in biomedical research. We explore the performance and scalability of HSCLO38 in specific use cases in handling large-scale genomic data in a biomedical knowledge graph.
Carrio Cordo, P.; Acheson, E.; Huang, Q.; Baudis, M.
Show abstract
Cancers arise from the accumulation of somatic genome mutations, which can be influenced by inherited genomic variants and external factors such as environmental or lifestyle-related exposure. Due to the heterogeneity of cancers, precise information about the genomic composition of germline and malignant tissues has to be correlated with morphological, clinical and extrinsic features to advance medical knowledge and treatment options. With global differences in cancer frequencies and disease types, geographic data is of importance to understand the interplay between genetic ancestry and environmental influence in cancer incidence, progression and treatment outcome. In this study, we analysed the current landscape of oncogenomic screening publications for geographic information content and quality, to address underrepresented study populations and thereby to fill prominent gaps in our understanding of interactions between somatic variations, population genetics and environmental factors in oncogenesis. We conclude that while the use of proxy derived geographic annotations can be useful for coarse-grained associations, the study of geo-correlated factors in cancer causation and progression will benefit from standardized geographic provenance annotations. Additionally, publication derived geographic provenance data allowed us to highlight stark inequality in the geographies of cancer genome profiling, with a near lack of sizeable studies from Africa and other large regions.
Mendez-Cruz, C.-F.; Diaz-Rodriguez, M.; Guadarrama-Garcia, F.; Lithgow-Serrano, O. W.; Gama-Castro, S.; Solano-Lira, H.; Rinaldi, F.; Collado-Vides, J.
Show abstract
The amount of published papers in biomedical research makes it rather impossible for a researcher to keep up to date. This is where machine processing of scientific publications could contribute to facilitate the access to knowledge. How to make use of text mining capabilities and still preserve the high quality of manual curation, is the challenge we focused on. Here we present the Lisen&Curate system designed to enable current and future NLP capabilities within a curation environment interface used in curation of literature on the regulation of transcription initiation in bacteria. The current version extracts regulatory interactions with the corresponding sentences for curators to confirm or reject accelerating their curation. It also uses an embedded metrics of sentence similarity offering the curator an alternative mechanism of navigating through semantically similar sentences within a given paper as well as across papers of a pre-defined corpus of publications pertinent to the task. We show results of the use of the system to curate literature in E. coli as well as literature in Salmonella. A major advantage of the system is to save as part of the curation work, the precise link for every curated piece of knowledge with the corresponding specific sentence(s) in the curated publication supporting it. We discuss future directions of this type of curation infrastructure.
Long, K.; Gravel-Pucillo, K.; Waldron, L.; Davis, S.; Oh, S.
Show abstract
Public omics repositories contain vast amounts of valuable data, but their metadata suffers from extreme heterogeneity, unstandardized terminologies, and quality issues that severely limit data reusability and cross-study integration. While prospective metadata standards exist, the majority of published omics data remain in non-standardized formats requiring retrospective curation. We performed comprehensive manual curation and harmonization of clinical metadata from 212,027 samples across 468 studies in two major repositories: curatedMetagenomicData (93 studies, 22,588 samples) and cBioPortal (375 studies, 189,438 samples). Through systematic ontology mapping, we consolidated redundant, dispersed information into much fewer harmonized columns, reduced unique values, and increased the completeness of major attributes. This curation process revealed common metadata quality issues, including typos, inconsistent terminologies, misplaced values, conflicting annotations, and inappropriately merged information across attributes. We document the challenges, decisions, and solutions encountered during large-scale metadata harmonization across two distinct omics domains. The harmonized metadata, accessible through the OmicsMLRepoR Bioconductor package, enables repository-wide queries and cross-study analyses previously challenging with heterogeneous metadata. Our experience provides practical guidance for similar curation efforts and demonstrates the value of investing in retrospective metadata improvement for existing public omics resources.
Li, Z.; Zhang, T.; Li, X.; Huang, J.; Xie, Z.; Gao, F.; Cai, H.; Sun, M.; Dai, M.; Liao, M.
Show abstract
Single-cell RNA sequencing (scRNA-seq) has dramatically advanced the understanding of cellular heterogeneity. While numerous marker gene databases are available for humans and mice, a lack of systematic resources for livestock and poultry species remains, limiting progress in functional genomics, immunology, and breeding.. To address this challenge, we developed AniMarkerDB (https://animarkerdb.bio), a comprehensive and curated database dedicated to marker genes and immune-related epitopes in economically animals, including chicken, pig, and duck. AniMarkerDB integrates 7,010 marker gene across 37 tissues and 846 cell types, together with 71,442 immune epitope records from IEDB. All entries undergo rigorous literature curation, manual validation, and multi-level quality control, with standardized nomenclature and annotation to ensure data consistency and reusability. The platform supports flexible queries by species, tissue, cell type, or gene. It offers analytical tools for cross-species comparison model organisms such as human and mouse, interactive single-cell atlas visualization, and user-defined cell type annotation. Additionally, AniMarkerDB provides dynamic visualizations and export options, enabling researchers to efficiently obtain large-scale marker and epitope data for downstream applications such as infectious disease research, vaccine target design, and comparative immunology. Looking ahead, AniMarkerDB will expand species coverage and incorporate additional modalities, including single-cell atlases from healthy and disease models, establishing itself as a comprehensive and authoritative platform for animal cell biology, disease modeling, and translational research. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=136 SRC="FIGDIR/small/682327v1_ufig1.gif" ALT="Figure 1"> View larger version (38K): org.highwire.dtl.DTLVardef@af7e7borg.highwire.dtl.DTLVardef@198d114org.highwire.dtl.DTLVardef@1c69ed2org.highwire.dtl.DTLVardef@e4fc43_HPS_FORMAT_FIGEXP M_FIG C_FIG
Caballero Perez, J.; de Blas Perez, C.; Leza Alvarez, F.; Caballero Contreras, J. M.
Show abstract
COVID-19 has had an unprecedented global impact in health and economy affecting millions of persons world-wide. To support and enable a collaborative response from the global research communities, we created a data collection for different public sources for anonymized patient clinical data, imaging datasets, molecular data as nucleotide and protein sequences for the SARS-CoV-2 virus, reports of count of cases and deaths per city/country, and other economic indicators in Databiology Lab (https://www.lab.databiology.net/) where researchers could access these data assets and use the hundreds of available open source bioinformatic applications to analyze them. These data assets are regularly updated and was used in a successful virtual 3-day hackathon organized by Databiology Ltd and Mindstream-AI where hundreds of attendees to work collaboratively to analyze these data collections.
Mai, Y.; lee, k.; Liu, Z.; Raja, K.; Higashi, M. K.; Ma, M.; Wang, T.; Ai, L.; Calay, E.; Oh, W.; Schadt, E.; wang, x.
Show abstract
The use of electronic health records (EHRs) holds the potential to enhance clinical trial activities. However, the identification of eligible patients within EHRs presents considerable challenges. We aimed to develop a pipeline for phenotyping eligibility criteria, enabling the identification of patients from EHRs with clinical characteristics that match those criteria. We utilized clinical trial eligibility criteria and patient EHRs from the Mount Sinai Database. The criteria and EHR data were normalized using national standard terminologies and in-house databases, facilitating computability and queryability. The pipeline employed rule-based pattern recognition and manual annotation. Our pipeline normalized 367 out of 640 unique eligibility criteria attributes, covering various medical conditions including non-small cell lung cancer, small cell lung cancer, prostate cancer, breast cancer, multiple myeloma, ulcerative colitis, Crohns disease, non-alcoholic steatohepatitis, and sickle cell anemia. 174 were encoded with standard terminologies and 193 were normalized using the in-house reference tables. The agreement between automated and manual normalization was high (Cohens Kappa = 0.82), and patient matching demonstrated a 0.94 F1 score. Our system has proven effective on EHRs from multiple institutions, showing broad applicability and promising improved clinical trial processes, leading to better patient selection, and enhanced clinical research outcomes.
PAN, F.; Zhang, Y.; Wang, J.; Liu, M.-C.; Sui, X.; Yue, H.; Zhang, J.
Show abstract
Infectious and immune-mediated diseases (IIDs) represent a broad and rapidly expanding biomedical literature domain in which scalable evidence extraction, disease ontology refinement, and interpretable knowledge integration are essential for biomedical discovery. We constructed an IID-specific biomedical knowledge graph (IID KG) from PubMed abstracts and PMC full-text articles by integrating nested named entity recognition, ontology-guided identifier assignment, full-text relation extraction, and relation-resolution strategies. A gold-standard corpus of 500 PubMed abstracts and 8 PMC full-text articles was manually annotated for nested biomedical entities across six entity types. The resulting models were applied to 30,128,068 PubMed abstracts and 1,385,500 IID-related PMC full-text articles. A unified IID ontology was developed from 411,341 disease terms using hierarchical text classification, large language model-based refinement, ontology cross-referencing, and expert review, yielding 179,657 confirmed MeSH mappings. The final IID KG contains approximately 1,837,513 unique entities and 16,295,390 unique relations across eight relation types. The resource was released publicly together with repurposing workflows, supporting ontology-aligned literature mining, disease mechanism analysis, and drug-repurposing hypothesis generation for IID research.
Sulieman, L.; Wu, P.; Denny, J.; Bastarache, L.
Show abstract
Researchers utilizing phenotypic data from diverse sources require matching of phenotypes to standard clinical vocabularies. Mapping phenotypes to vocabulary can be difficult, as existing tools are often incomplete, can be difficult to access, and can be cumbersome to use, especially for non-experts. We created WikiMedMap as a simple tool that leverages Wikipedia and maps phenotype strings to standard clinical vocabularies. We assessed WikiMedMap by mapping phenotype strings from questionnaires in the UK Biobank and from Mendelian diseases in Online Mendelian Inheritance in Man (OMIM) database to eight vocabularies: International Classification of Diseases, Ninth Revision (ICD-9), ICD-10, ICD-O, Medical Subject Headings (MeSH), OMIM, Disease Database, and MedlinePlus. WikiMedMap outperformed conventional mapping tools in finding potential matches for phenotype strings. We envision WikiMedMap as a technique that complements existing and established tools to map strings to clinical vocabularies that usually do not coexist in one source.
Nandani, T.; Ott, B. P.; Balaratnam, P.; Archer, S. L.; Durbin, J.; Hindmarch, C. C. T.
Show abstract
Pulmonary hypertension (PH) is a vasculopathy that results in elevated mean pulmonary arterial pressures over 20mmHg. Despite significant advances in research, PH still has a high mortality rate, and there is currently no cure for the disease. As with all biomedical fields, PH researchers have embraced the power of next generation technologies such as microarrays and RNA sequencing. Most of these data can be found on public repositories, which is usually a requirement for publication. While these repositories are rich sources of data, they require intermediate to advanced bioinformatics skills to access, download, and make these data useful. Here we present Pulmonary Hypertension Engine for Linked Experiments (PHELEX), which represents a comprehensive catalogue of all RNA sequencing data related to PH that is currently available on the Gene Expression Omnibus (GEO), hosted by the US National Centre for Biotechnology Information (NCBI). We identified 2,278 bulk RNA sequencing samples from human, mouse and rat, and built a searchable tool based on the metadata that is associated with each sample. PHELEX is a functional tool that allows selected studies to be highlighted, and parsed through Confidence, an analysis tool we have created, which will model the data based on user-defined classifiers, perform differential gene expression and pathway analysis, and present these data using standard graphics, and text-file results. PHELEX also allows PH researchers to cross-cut between discrete studies, facilitating de novo understanding of these data. As a robust searchable repository of genomic data, we hope that PHELEX will accelerate PH innovation and discovery, by allowing researchers to mine existing genomic data and thus better understand the molecular signatures that underpin PH.
Alganem, K.; Shukla, R.; Eby, H.; Abel, M.; Zhang, X.; Mcintyre, W. B.; Lee, J.; Au-Yeung, C.; Asgariroozbehani, R.; Panda, R.; O'Donovan, S. M.; Funk, A.; Hahn, M.; Meller, J.; McCullumsmith, R.
Show abstract
BackgroundIn silico data exploration is a key first step of exploring a research question. There are many publicly available databases and tools that offer appealing features to help with such a task. However, many applications lack exposure or are constrained with unfriendly or outdated user interfaces. Thus, it follows that there are many resources that are relevant to investigation of medical disorders that are underutilized. ResultsWe developed an R Shiny web application, called Kaleidoscope, to address this challenge. The application offers access to several omics databases and tools to let users explore research questions in silico. The application is designed to be user- friendly with a unified user interface, while also scalable by offering the option of uploading user-defined datasets. We demonstrate the application features with a starting query of a single gene (Disrupted in schizophrenia 1, DISC1) to assess its protein-protein interactions network. We then explore expression levels of the gene network across tissues and cell types in the brain, as well as across 34 schizophrenia versus control differential gene expression datasets. ConclusionKaleidoscope provides easy access to several databases and tools under a unified user interface to explore research questions in silico. The web application is open-source and freely available at https://kalganem.shinyapps.io/Kaleidoscope/. This application streamlines the process of in silico data exploration for users and expands the efficient use of these tools to stakeholders without specific bioinformatics expertise.
Lin, Y.-J.; Menon, A. S.; Hu, Z.; Brenner, S. E.
Show abstract
BackgroundVariant interpretation is essential for identifying patients disease-causing genetic variants amongst the millions detected in their genomes. Hundreds of Variant Impact Predictors (VIPs), also known as Variant Effect Predictors (VEPs), have been developed for this purpose, with a variety of methodologies and goals. To facilitate the exploration of available VIP options, we have created the Variant Impact Predictor database (VIPdb). ResultsThe Variant Impact Predictor database (VIPdb) version 2 presents a collection of VIPs developed over the past 25 years, summarizing their characteristics, ClinGen calibrated scores, CAGI assessment results, publication details, access information, and citation patterns. We previously summarized 217 VIPs and their features in VIPdb in 2019. Building upon this foundation, we identified and categorized an additional 186 VIPs, resulting in a total of 403 VIPs in VIPdb version 2. The majority of the VIPs have the capacity to predict the impacts of single nucleotide variants and nonsynonymous variants. More VIPs tailored to predict the impacts of insertions and deletions have been developed since the 2010s. In contrast, relatively few VIPs are dedicated to the prediction of splicing, structural, synonymous, and regulatory variants. The increasing rate of citations to VIPs reflects the ongoing growth in their use, and the evolving trends in citations reveal development in the field and individual methods. ConclusionsVIPdb version 2 summarizes 403 VIPs and their features, potentially facilitating VIP exploration for various variant interpretation applications. AvailabilityVIPdb version 2 is available at https://genomeinterpretation.org/vipdb