Database
◐ Oxford University Press (OUP)
Preprints posted in the last 90 days, ranked by how well they match Database's content profile, based on 61 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.
Cook, N.; Boulais-Richard, J.; Zeng, Y.; Yang, C.; Budde, J.; Taliun, D.; Gagliano Taliun, S. A.; Cruchaga, C.; Belloy, M. E.
Show abstract
Summary: The X chromosome comprises approximately 5% of the human genome and encodes over 800 protein-coding genes, many of which exhibit sex-differentiated expression patterns due to escape from X chromosome inactivation (XCI) mechanisms. Despite its relevance to sex differences in complex traits, the X chromosome is routinely excluded from genome-wide association studies due to analytical challenges, and when analyzed, the impact of escape from XCI or sex is limitedly explored. No dedicated, publicly accessible browser for X chromosome-wide association study (XWAS) summary statistics currently exists, creating a barrier to systematic investigation of X-linked contributions to human traits. Here, we present geneXplore, an interactive web browser based on the PheWeb2 implementation, tailored for XWAS summary statistics across 1,944 phenotypes while distinguishing random XCI (rXCI), escape from XCI (eXCI), and sex-stratified analyses. Users can explore results via interactive plots (Manhattan and Miami, PheWAS and LocusZoom), searchable tables and access to cross-database lookup, with full summary statistics available for download. Availability and Implementation: geneXplore is freely available at https://genexplore.wustl.edu/ with no registration required and will be maintained for a minimum of two years following publication. Source code is available at https://github.com/Belloy-Lab/geneXplore_XWAS_Browser under an MIT license.
Vand, K.; Badia, N.; Khomtchouk, B.; Janga, S. C.
Show abstract
Cardiovascular genomics is producing rapidly expanding genetic, molecular, phenotypic and clinical data, yet relevant evidence remains fragmented across resources and difficult to translate into actionable biological and ultimately translational knowledge. HeartBioPortal (HBP) is a browser-based cardiovascular knowledge environment that was developed to address this problem by organizing omics, variant, phenotype and clinical evidence centered around gene queries. Here we describe HBP 3.0, a major update that expands both the data architecture and interpretive interface. This update introduces DataHub, a reproducible data-engineering layer for source ingestion, standardization, variant-centered aggregation, provenance tracking and compact serving artifacts. The release integrates cardiovascular clinical practice guideline context through a graph-backed clinical knowledge layer; incorporates cardiovascular summary statistics from the Million Veteran Program and public aggregate resources; expands source-preserving population frequency, variant annotation and structural-variant; and adds gene profile, drug-discovery and protein-context layers. HBP 3.0 incorporates 594.3 million allele-frequency observations across 18.1 million rsIDs, 3.04 million exon-enriched structural-variant records, 66.9 thousand protein isoforms with 3.26 million non-exon protein feature annotations, 17,128 gene-drug records, and a clinical guideline knowledge graph with 42,895 entities and 106,304 relationships. The redesigned gene dossier view combines phenotype filtering, annotation composition, persistent selected-detail panels and exportable chart data in one workflow. HBP 3.0 is designed to help cardiovascular and eventually cardiometabolic researchers move from a genetic or genomic signal to biological knowledge and potentially clinical and therapeutic context while preserving source provenance and interpretive boundaries. Database URL: https://www.heartbioportal.com/
Gardeux, V.; Carsanaro, S.; Chen, W. J.; David, F. P. A.; Goutte-Gattat, D.; Hilton, J. A.; Lubiana, T.; Patel, N.; Raymor, B.; Zucchi, I.; Deplancke, B.; Ernst, C.; Osumi-Sutherland, D.; Robinson-Rechavi, M.; Sternberg, P. W.; Bastian, F. B.
Show abstract
The rapid accumulation of single-cell RNA-Seq (scRNA-seq) data across multiple repositories presents major challenges for data accessibility, integration, and reproducibility. While primary repositories provide raw data, they rarely include structured cell-type annotations or descriptions of analytical workflows, limiting the ability to reuse and integrate datasets in a FAIR (Findable, Accessible, Interoperable, Reusable) manner. Here we present scFAIR, a consortium of single-cell data resources that has developed a unified metadata schema and common curation framework to improve the FAIRness of scRNA-seq data. Building on and extending the CZ CELLxGENE Discover metadata schema, the scFAIR consortium has been instrumental in driving key schema improvements, including the expansion of supported organisms, richer biological context, and structured reporting of computational workflows. To provide unified access to decentralized datasets, the consortium developed the sc-fair.org portal, which currently aggregates 2,346 datasets across partner resources through ontology-aware semantic search. We demonstrate the practical value of FAIR-compliant datasets through a cross-species validation between human and mouse Allen Brain Atlases, showing that standardized ontology annotations enable reliable annotation transfer across species, with 90% of neuronal clusters receiving an exact or equivalent label. Together, the scFAIR schema, validator, and portal constitute a community-driven framework that advances single-cell data standardization and lays the foundation for reproducible, large-scale integration of single-cell datasets.
Nehra, N.; Swami, R.; Dadi, D.; Mishra, R.; Sharma, U.; Verma, P.; Sen, M.; Dhruw, N. K.; Jha, A. K.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWGetting clinical data from different sources to "talk" to each other within the OMOP Common Data Model (CDM) is arguably the most tedious part of multi-center research. While this integration is essential, the transformation process is frequently a manual grind, requiring a rare overlap of deep clinical knowledge and technical expertise. In this paper, we present a framework designed to alleviate some of the burden on the researcher by automating data harmonization through two distinct steps: structural schema mapping and terminological standardization. For the structural piece, we moved away from "black box" logic in favor of a stateful workflow managed by large language models (LLMs) and directed acyclic graphs. By profiling EHR data at the source, our system generates context-aware dictionaries that offer ranked mapping suggestions alongside confidence scores. While our benchmarking showed a 97.5% agreement rate at the schema level and an 84% agreement rate at the value level when compared with human experts, the system appears most effective when treated as a "co-pilot" rather than a total replacement for human oversight. To handle value-level standardization, we implemented a hybrid search strategy that pairs the semantic depth of SapBERT embeddings with the literal precision of fuzzy string matching. By using FAISS for rapid similarity retrieval, the engine attempts to resolve messy or "noisy" clinical descriptions to standard OMOP concepts. This approach seems particularly promising for handling the non-standardized labels that often plague smaller, local datasets. Ultimately, our results suggest that this guided approach can shift the timeline for OHDSI-compliant warehousing from weeks of manual curation to a more manageable and scalable pipeline, potentially lowering the barrier to entry for smaller research teams.
Glen, A. K.; Witherington, D.; Leslie, T.; Baumgartner, A.; Fernando, A.; Vemuri, B.; Nahman, O.; Glusman, G.; Hood, L.; Pflieger, L.; Rappaport, N.
Show abstract
Existing general-purpose biomedical knowledge graphs tend to focus on disease mechanisms and drug repurposing, leaving multiomic and wellness-relevant content underrepresented. KRAKEN (Knowledge Research & Analysis Kit for Evidence Networks) addresses this gap by integrating existing graphs (including Translator KG Open, RTX-KG2, and ROBOKOP) with specialized sources such as RefMet, LIPID MAPS, NIH Common Data Elements, Polygenic Score Catalog, and derived wellness measures including biological age and biological BMI. The resulting graph spans ~15M nodes and ~113M edges across 62 entity types. KRAKEN adopts the Biolink Model as its semantic layer, ensuring compatibility with standardized resources emerging from the NIH NCATS Biomedical Data Translator program. A lightweight, modular build system rebuilds the full graph (including entity resolution), with peak memory consumption under 48 GB, and supports flexible inclusion or exclusion of sources, allowing the user to scope the graph to a domain of interest. Built-in analytical tools include multi-hop reasoning, subgraph extraction, text, vector and hybrid entity search, and enrichment analyses, all accessible through an interactive web interface, a REST API, and a Model Context Protocol server, the last enabling direct consumption by agentic and LLM-based systems. KRAKEN is freely available at https://app.krakenkg.com.
Golnari, P.; Prantzalos, K.; Upadhyaya, D. P.; Buchhalter, J.; Sahoo, S. S.
Show abstract
Dravet syndrome (DS) is a severe developmental and epileptic encephalopathy whose clinical and research representation requires integration of heterogeneous knowledge spanning seizures, development, behavior, SUDEP/autonomic risk, genetics, comorbidities, electrophysiology, pharmacology, and drug responsiveness. We report the development of a DS-focused ontology created by expert-guided specialization of a previously published epilepsy ontology. Scope expansion was defined through a scientific advisory board, structured review meetings, and iterative ontology curation in OWL. The resulting resource reorganized DS content across nine major domains and expanded the publicly released ontology from the pre-extension baseline to the current BioPortal version. Beyond structural growth, the ontology was assessed through expert-guided curation and downstream task-based reuse, including two published ontology-enabled LLM studies and an ongoing ontology-derived DS knowledge graph and AI assistant platform. These results suggest that disease-focused ontology specialization can provide durable infrastructure for DS data harmonization, knowledge representation, and AI-enabled translational informatics.
Zhou, Y.; Huang, F.; Zhao, Y.
Show abstract
Tuberculosis remains a major global public health threat. While whole-genome sequencing has transformed our understanding of the causative agent, Mycobacterium tuberculosis (MTB), existing genomic databases are highly fragmented and often underrepresent structural variations (SVs). Furthermore, critical population-genetic statistics are rarely integrated with phylogenetic and geographic context, forcing researchers to reconcile separate datasets manually. To address this gap, we developed TBpop (https://tbpop.chinacdc.cn), an open-access, integrated population genomics portal. TBpop is built from 420 clinical MTB isolates selected from the first national drug resistance baseline survey in China. The portal integrates isolate metadata, pangenome categories, SNPs, SVs, IS6110 insertion sites, strain phylogeny, and gene-level statistics, and provides three interactive explorer modules: the Population Explorer, the Statistics Explorer, and the Variation Explorer. Additionally, a User Analysis module allows researchers to run population genetic workflows on their own alignments. TBpop provides an integrated platform for exploring genome plasticity, signatures of positive selection, and conservation patterns of functionally important genes in MTB.
Sun, S.; Wang, H.; Mathe, E. A.; Zhu, Q.
Show abstract
Rare diseases (RD) impact over 30 million individuals in the United States, yet fewer than 5% of the identified conditions have FDA-approved treatments. Progress in RD research is hindered by small patient cohorts, biological heterogeneity, and the fragmented, inconsistently annotated publicly available omics data, which limits integrative analysis and translational discovery. Here, we present RD-OMICS, a data inventory with integrated and structured RD omics data from Gene Expression Omnibus (GEO), in the form of a knowledge graph. We developed a metadata harmonization pipeline that combines rule-based mapping and large language model (LLM)-assisted semantic categorization. The graph-based data model was defined to integrate different types of data including disease conditions, experiments, samples, platforms, projects, and publications into a centralized inventory graph. In this preliminary study, 11,049 GEO series for 126 rare diseases were processed and integrated into RD-OMICS, which includes 375,930 individual biospecimen samples, 1,578 sequencing and array platforms, 10,938 biological projects. Case studies demonstrate the use of RD-OMICS in supporting rare disease research, omics cohort construction, and transcriptome-based drug repurposing for amyotrophic lateral sclerosis (ALS). RD-OMICS provides a scalable foundation for transforming fragmented omics data into a structured, harmonized and interoperable resource, facilitating therapeutic development and other translational discoveries in rare diseases.
Lieutaud, P.; McLaughlin, j.; Hendrickson, R. C.; David, R.; Parkinson, H.; Lefkowitz, E.; Dempsey, D.; Coutard, B.
Show abstract
The International Committee on Taxonomy of Viruses (ICTV) is responsible for developing and maintaining a universal virus taxonomy. As the reference framework for organising the viral world, it is essential for virology and related fields. Despite its widespread use in research and public health, programmatic access to ICTV taxonomy has remained limited, posing challenges for integration, versioning, and interoperability across databases and bioinformatics resources requiring up-to-date virus taxonomy. To address this, we developed a public and sustainable solution leveraging ontology-based APIs. Successive ICTV Master Species List (MSL) releases were transformed into a structured ontology and deployed as a unified representation through the Ontology Lookup Service (OLS). The framework also provides ICTV-NCBI mappings and helper libraries for integration into downstream systems. This enables, for the first time, public programmatic retrieval of current and historical virological taxon names, taxonomic relationships, metadata, and persistent identifiers through stable endpoints. More broadly, this work illustrates a general strategy for transforming structured biological datasets into semantically enriched graph resources exposed through scalable public APIs. These developments enhance interoperability, reduce manual curation, and support FAIR-aligned taxonomic data management in virology and pandemic preparedness. Key pointsO_LIICTV provides the official taxonomy for classifying viruses and naming virus taxa, but lacks standardised programmatic access. C_LIO_LITransforming ICTV data into an ontology enables semantic, machine-actionable access across releases via ontology-based APIs. C_LIO_LIICTV-NCBI mappings support interoperability across bioinformatics resources. C_LIO_LIThe framework enables programmatic resolution of current and historical viral taxa. C_LIO_LIThis approach provides a reusable model for exposing biological datasets through public APIs. C_LI
Murad, A. B.
Show abstract
BackgroundReuse of archived transcriptomic data underpins a large and growing share of published genomics. Because differences in upstream processing confound cross-study comparison, uniform reprocessing compendia -- recount3, ARCHS4, DEE2, refine.bio, Expression Atlas -- are widely treated as the remedy, and their availability is routinely assumed at the point of study design. Whether that remedy is actually obtainable for the population of published disease RNA-seq has not been measured. Prior audits have characterised metadata completeness and deposition rates, but none has quantified, across the published population, what fraction of studies can be uniformly reprocessed or where in the path from publication to comparable counts that capability is lost. ResultsWe enumerated 1,124 MeSH disease descriptors exhaustively, retrieved 16,820 human RNA-seq series from the Gene Expression Omnibus, and audited the 3,631 bulk, Illumina-platform series of at least 25 samples under three independently pre-registered, tool-enforced analysis plans. Raw reads were publicly available for 94.1% of series under a dual-route evidence standard, but only 46.7% appeared in any uniform reprocessing compendium (bounds 46.7-57.2%) and only 27.3% were usably covered at a 90% run threshold (bounds 27.3-38.8%). Of the 3,418 series whose reads are public, 991 were usably covered, leaving 71.0% of read-public series reprocessed by nothing usable. Presence overstated usability: DEE2 was present for 30.2% of series but usable for 3.4%. Design attrition was independent and severe -- 24.7% met bulk primary-tissue case-control criteria, 7.4% additionally reached a minimum replication threshold counted on sample accessions, and 4.1% did so counted on distinct donors. Among the 199 series where donor identity resolves, 4.1% pass the replication criterion on donors against 15.1% on accessions, a 3.73-fold difference; across the census frame, accessions exceeded distinct donors by 2.65-fold (Manski bounds 1.09-10.77, Imbens-Manski 95% CI 1.07-11.88). Independently, 36.4% of series-to-disease attributions produced by a conventional keyword query were refuted by the curated MeSH headings of the series own linked publication. ConclusionsUniform reanalysis of published human disease RNA-seq is unavailable for most studies in the population audited, and the binding constraint is usable coverage rather than deposition of raw reads. The loss occurs at several independent layers with different remedies, and the coverage layer -- unlike the others -- is one that resource maintainers can act on. Automated retrieval further over-counts eligible studies, both by admitting designs outside scope and by assigning studies to diseases their publications do not support.
Srinivasan, S.; Chande, A.
Show abstract
Post-transcriptional chemical modifications of RNA, collectively termed the epitranscriptome, have emerged as critical regulatory layers governing viral replication, pathogenicity, and host-virus interactions. Despite the rapid accumulation of experimental data on viral RNA modifications, no dedicated, freely accessible resource existed for systematically cataloguing these sites across diverse viral species. Here we present ViralEpiBase, a manually curated database of epitranscriptomic modification sites identified in viral RNA genomes and virus-encoded transcripts at single-nucleotide resolution. ViralEpiBase currently integrates seven chemically distinct RNA modification types: N6-methyladenosine (m6A), N1-methyladenosine (m1A), pseudouridine ({Psi}), 5-methylcytosine (m5C), 2'-O-methylation (2'OMe), inosine and N4-acetylcytidine (ac4C); across 12 viral species encompassing both DNA and RNA viruses of clinical and biological significance. Each entry is linked to its primary literature source or deposited dataset and is retrievable by modification type, genomic coordinates, or viral taxonomy. The database is freely accessible through an intuitive web interface and is updated continuously as new experimental evidence becomes available. ViralEpiBase thus provides the first unified platform dedicated exclusively to viral epitranscriptomics and is designed to facilitate mechanistic investigation of RNA modification functions in viral biology.
Tindall, C.; Long, R. A.; Naughton, B.; Mapes, B. M.; Vismer, D.; Skinner, H. G.; Malenfant, J.; Maurya, M. R.; Nalls, M. A.; Ramachandran, S.; Nguyen, T.; Peters, M. A.; Scheuermann, R. H.
Show abstract
SysBio FAIRplex is a Common Fund Venture Program that catalogs and indexes data from the Accelerating Medicines Partnership(R) (AMP(R)) Program through a federated model in which data hosts retain custody of their datasets. The central piece of this work is the SysBio Common Data Model (SysBio CDM). AMP is a precompetitive public-private partnership started in 2014 that unites the resources of NIH and private partners to improve our understanding of disease pathways and transform current models for developing new treatments by: - identifying new targets, biomarkers, and development paradigms; - developing leading-edge tools and technologies; - collecting large-scale datasets and supporting analytics for open analysis by the public; and - generating consensus platforms and procedures. A multidisciplinary Task Force was chartered to design the SysBio CDM by extending the Observational Medical Outcomes Partnership (OMOP) Common Data Model into the -omics domain. The Task Force produced a Minimum Viable Product comprising nine OMOP tables; four extension tables for assay and file metadata; and a Common Data Element (CDE) Registry to specify field semantics. This manuscript describes the deliverable: the underlying design choices, the criteria applied in selecting and constructing the extension tables, how the extended model supports multimodal data integration across AMP projects, and what further work to support additional -omics modalities would entail. As an auxiliary methodology, the paper also describes the AI-assisted CDE harmonization workflow used to populate the model.
Medina-Ortiz, D.; Olivera-Nappa, A.; Lienqueo, M. E.; Opazo, R.; Romero, J.
Show abstract
Bacteriophage lytic enzymes and depolymerases are relevant to phage biology, antimicrobial development, and protein engineering, but their sequence and annotation data remain dispersed across general databases, specialized resources, genome-centred collections, and prediction-oriented datasets. We present PhageLysData, an evidence-aware and AI-ready resource constructed through reproducible multisource integration, provenance tracking, and exact-sequence consolidation. The release integrates 807,366 source observations from seven primary resources into 759,105 unique exact-sequence entities, comprising an evidence-supported Core of 11,867 entities, a Prediction Extension of 745,092 prediction-only candidates, and 2,146 Context entities retained for provenance and reference. This architecture preserves broad sequence-space coverage while maintaining a clear distinction between non-predictive and prediction-derived support. Core entities are enriched with harmonized biological annotations, physicochemical properties, independent InterProScan-derived functional annotations, mapped PDB and AlphaFold DB structural assets, and reusable numerical representations. For 11,259 eligible Core sequences, PhageLysData provides embeddings from 11 protein language models together with one-hot encoding under a common representation contract. Release-facing examples demonstrate latent-space exploration, unsupervised clustering, supervised classification, and evidence-aware candidate retrieval without defining a universal predictive benchmark. PhageLysData provides a traceable, versioned, and computationally accessible foundation for protein retrieval, comparative analysis, task-specific dataset construction, and machine-learning applications involving phage lytic enzymes and depolymerases.
NANDI, S.; Sundararajan, Z.; Subirana-Granes, M.; Espinosa, J. M.; Pividori, M.; Sullivan, K. D.; Galbraith, M. D.; Costello, J.
Show abstract
Down syndrome, caused by trisomy 21, increases the risk of diverse co-occurring conditions. With more than 34,000 related publications indexed in PubMed as of early 2026, keeping pace with this expanding literature is challenging. While general-purpose large language models are widely used for information retrieval, they often rely on broad training data rather than specific evidence. Retrieval-augmented generation (RAG) improves rigor and reliability of responses by linking model outputs to source texts. In research, source texts are peer-reviewed articles. Standard implementations treat all manuscript sections equally, allowing background text to rank as highly as experimental results. To focus model outputs on experimentally supported responses, we developed the T21 Research Assistant, a section-aware RAG system that prioritizes Results sections to ground responses in primary experimental evidence. The system draws exclusively from 1,789 open-access Down syndrome publications from PubMed Central, including 327 NIH INCLUDE-funded studies, and uses a multistage pipeline for query validation, retrieval, reranking, synthesis, and citation verification. Built on NVIDIA Nemotron models, it generates structured, cited responses. Evaluation using expert-curated questions demonstrated strong performance, achieving a BERTScore F1 of 0.712 and recall of 0.758, comparable to or exceeding leading proprietary and open-source models. T21 Research Assistant is available at: https://bioinformatics.cuanschutz.edu/t21-res-assi/
Chi, L. A.; Ytreberg, F. M.
Show abstract
Protein-peptide interactions are central to cellular signaling and to a growing class of peptide therapeutics, yet the datasets used to develop and benchmark computational methods for protein-peptide modeling remain poorly standardized. Available databases prioritize comprehensive coverage but require task-specific curation, while published benchmarks are typically distributed as static collections built with heterogeneous curation, quality-filtering, redundancy-reduction, and sampling strategies, limiting reproducibility and cross-study comparison. We present PepXPro, a modular framework that transforms publicly available protein-peptide structure-affinity resources into curated datasets and reproducible benchmark collections generated under user-defined criteria. PepXPro is organized into three components: Scrape, for deterministic curation of protein-peptide complex entries from public resources; GenSample, for constructing configurable subsets under explicit quality, redundancy, and sampling constraints; and Benchmark, for evaluating candidate subsets and selecting a nonredundant, representative, general-purpose benchmark for distribution. Starting from PDBbind and complementary resources, the curation pipeline yields a pool of proteinpeptide complex entries that retains chemically complex cases, including disulfide- linked cyclic peptides, which are commonly excluded from existing benchmarks. We release PepXPro Benchmark v1, a benchmark comprising 70 non-redundant protein- peptide complexes with experimentally determined structures and binding affinities. The underlying framework provides an extensible foundation for reproducible protein- peptide benchmark construction.
Xuan, H.; Pasupuleti, R.; Liu, B.; Sun, H.; Zhang, J.; Yao, Z.; Zhong, C.
Show abstract
Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale. The lack of a comprehensive and up-to-date catalog of bioinformatics resources hinders efforts toward automated biomedical knowledge extraction and streamlined data analysis. Here we present SNAIL, a hybrid named entity recognition framework designed to automatically identify bioinformatics software and database (SW/DB) names from biomedical texts. SNAIL integrates complementary lexical and semantic modeling strategies. The lexical component captures orthographic patterns and contextual cues characteristic of SW/DB names, while the semantic component leverages contextual embeddings generated by transformer-based language models such as SciBERT, combined with an explicit token-masking strategy to enhance entity-focused representations. A large training corpus was constructed automatically through a hybrid pipeline that integrates citation-hinted extraction with large language model-assisted distillation. Evaluation on two independent benchmark datasets and real-world research articles demonstrates that SNAIL substantially outperforms existing approaches, including domain-specific methods such as bioNerDS2 and general-purpose large language models such as ChatGPT, Gemini, Grok and Claude. Applying SNAIL to large-scale literature analysis further reveals distinct journal-level preferences across bioinformatics subfields. These results demonstrate that SNAIL provides an accurate and scalable solution for identifying bioinformatics resources in scientific texts and enables systematic meta-analysis of tool usage and research trends.
Pavlidis, P.; Mancarci, B. O.; Maximo, A.; Yan, C.; Schwartz, R. A.
Show abstract
We describe an automated software tool to accomplish data curation tasks previously performed by humans for the Gemma genomics data re-analysis resource. Gemma is a hand-curated database of reprocessed transcriptomic studies, currently covering over 23,000 human, mouse and rat data sets largely drawn from the Gene Expression Omnibus (GEO). We developed a pipeline that uses both traditional (mechanical) and large-language models to produce detailed ontology-anchored, sample- and experiment-level annotations in accordance with our established curation guidelines. In this report, we describe benchmarking the pipeline and investigations aimed at evaluating readiness of the v1.1 Gemma curation agent for production use. Overall, performance is near that of human curators, at approximately 1/20th the cost and at least 100 times the speed. We also present preliminary exploration of triage methods for identifying agent curations that are more likely to contain errors, and thus can be forwarded for human review. We discuss the potential place of such curation approaches in bioinformatics ecosystems. Besides the software, our deliverables include the benchmark set of 500 studies and an evaluation framework that can be used to further develop the pipeline or compare to other approaches.
Nguyen, T. T. D.; Nguyen-Phuong, T.; Nguyen, Q.-H.; Abbasi, A. M.; Le Phan, H.-D.; Nguyen, L. B.-A.; Phan, N.-T.; Curabaz, N. N.; Hauser, A. S.; Tanoli, Z.; Nguyen, D. T.; Kooistra, A. J.
Show abstract
Biomedical knowledge evolves rapidly, yet most disease-centered knowledge graphs remain unchanged after publication. We present PrimeKG-Plus, an extension of PrimeKG that updates all 20 original data resources to releases available as of Dec 2025 and incorporates additional resources, including OpenTargets, RepurposeDrugs, and nSIDES. The graph is further expanded with relations extracted from 637 PubMed abstracts and PMC full-text articles using a large-language-model-assisted curation workflow. Extracted relations were refined through normalization, UMLS-based synonym mapping, SapBERT embedding based similarity ranking, and human expert review to ensure data quality. This literature-driven expansion focuses on rare neurological disorders, including Canavan disease, Niemann-Pick disease type C, Tay-Sachs disease and Batten disease. Network topology analyses indicate that the expanded graph improves indirect drug-disease connectivity of 3-6 hops through the graph while adding 447,288 Drug[->]Protein[->]Disease paths linking drug-disease pairs that were previously unreachable. Temporal FDA validation captured 55 new molecular entities after the June 2021 PrimeKG data cut-off as drug nodes, 46 absent from the original PrimeKG. By integrating newly incorporated and updated resources with systematically curated literature-derived associations, PrimeKG-Plus provides an up-to-date knowledge graph for network-based drug repurposing and precision medicine applications. Data and code are publicly available.
McLaughlin, J.; Puig-Barbe, A.; Ibrahim, A.; Pava, D.; Pendlington, Z. M.; Matentzoglu, N.; Sollis, E.; Foreman, A.; Wilson, R.; Lopez Gomez, F.; Harris, L.; Adeleye, Y.; Kaur, S.; Meldal, B.; Smedley, D.; Parkinson, H.
Show abstract
Researchers increasingly need to explore hypotheses that span multimodal data across different scales, organisms, and domains. In practice, this requires connecting knowledge across fragmented databases with incompatible APIs and heterogeneous annotation practices. Large language model (LLM) agents can automate this data integration process, but grounding LLM agent outputs in scientifically correct sources of truth remains a significant challenge. Here we describe our deployment of a novel AI semantics workflow using LLM agents to enable scalable data integration, grounded in biological knowledge in the form of ontologies. Our workflow comprises (1) a multi-agent system curating scientific knowledge across ontologies using the Ontology Lookup Service (OLS) as grounding; (2) an LLM embedding service to enable interoperability between scientific databases by mapping ontology terms; and (3) GrEBI, a knowledge graph and Model Context Protocol (MCP) server enabling LLM agents to conduct cross-cutting, multi-omic biomedical queries. O_FIG O_LINKSMALLFIG WIDTH=187 HEIGHT=200 SRC="FIGDIR/small/742514v1_ufig1.gif" ALT="Figure 1"> View larger version (39K): org.highwire.dtl.DTLVardef@1fc525aorg.highwire.dtl.DTLVardef@82c4f7org.highwire.dtl.DTLVardef@15173fdorg.highwire.dtl.DTLVardef@961f1c_HPS_FORMAT_FIGEXP M_FIG C_FIG
Baker, M.; Bett, K.; Vargas, A.; Jin, L.
Show abstract
Structural variants (SVs) are large-scale genomic variants, which can disrupt important functional and regulatory elements, leading to genomic disorders in humans and playing important roles in domestication, disease resistance, and traits in plants. SVs are generated across populations of individuals and used for association studies, consisting of large datasets with thousands of genomic loci. Visualization of these SVs aids in understanding their genomic distribution, identifying patterns across affected or phenotypic groups, and assessing their proximity to other genomic regions of interest. A variety of tools exist for visualizing SVs, including linear genome browsers and graph-based methods; however, many do not offer intuitive or scalable representations of SVs across large populations. To address this, we present SVPopEx, an interactive tool for population-wide visualization and exploration of SVs. SVPopEx provides a unique and intuitive representation for insertions, deletions, inversions, duplications, and translocations in a linear genome-style browser. Novel features were developed to support comparisons across genomes within user-defined regions, including rendering SVs based on one or more samples and visualizing haplotypes. Use of the tool is demonstrated with SV datasets from Schistosoma mansoni and Lens culinaris. A task-based evaluation was conducted using SVPopEx and two other linear genome browsers, which demonstrated that SVPopEx excelled in (1) providing a clear representation of the SVs present and (2) supporting comparisons across genomes.