Database
◐ Oxford University Press (OUP)
Preprints posted in the last 30 days, ranked by how well they match Database's content profile, based on 61 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.
Nehra, N.; Swami, R.; Dadi, D.; Mishra, R.; Sharma, U.; Verma, P.; Sen, M.; Dhruw, N. K.; Jha, A. K.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWGetting clinical data from different sources to "talk" to each other within the OMOP Common Data Model (CDM) is arguably the most tedious part of multi-center research. While this integration is essential, the transformation process is frequently a manual grind, requiring a rare overlap of deep clinical knowledge and technical expertise. In this paper, we present a framework designed to alleviate some of the burden on the researcher by automating data harmonization through two distinct steps: structural schema mapping and terminological standardization. For the structural piece, we moved away from "black box" logic in favor of a stateful workflow managed by large language models (LLMs) and directed acyclic graphs. By profiling EHR data at the source, our system generates context-aware dictionaries that offer ranked mapping suggestions alongside confidence scores. While our benchmarking showed a 97.5% agreement rate at the schema level and an 84% agreement rate at the value level when compared with human experts, the system appears most effective when treated as a "co-pilot" rather than a total replacement for human oversight. To handle value-level standardization, we implemented a hybrid search strategy that pairs the semantic depth of SapBERT embeddings with the literal precision of fuzzy string matching. By using FAISS for rapid similarity retrieval, the engine attempts to resolve messy or "noisy" clinical descriptions to standard OMOP concepts. This approach seems particularly promising for handling the non-standardized labels that often plague smaller, local datasets. Ultimately, our results suggest that this guided approach can shift the timeline for OHDSI-compliant warehousing from weeks of manual curation to a more manageable and scalable pipeline, potentially lowering the barrier to entry for smaller research teams.
Glen, A. K.; Witherington, D.; Leslie, T.; Baumgartner, A.; Fernando, A.; Vemuri, B.; Nahman, O.; Glusman, G.; Hood, L.; Pflieger, L.; Rappaport, N.
Show abstract
Existing general-purpose biomedical knowledge graphs tend to focus on disease mechanisms and drug repurposing, leaving multiomic and wellness-relevant content underrepresented. KRAKEN (Knowledge Research & Analysis Kit for Evidence Networks) addresses this gap by integrating existing graphs (including Translator KG Open, RTX-KG2, and ROBOKOP) with specialized sources such as RefMet, LIPID MAPS, NIH Common Data Elements, Polygenic Score Catalog, and derived wellness measures including biological age and biological BMI. The resulting graph spans ~15M nodes and ~113M edges across 62 entity types. KRAKEN adopts the Biolink Model as its semantic layer, ensuring compatibility with standardized resources emerging from the NIH NCATS Biomedical Data Translator program. A lightweight, modular build system rebuilds the full graph (including entity resolution), with peak memory consumption under 48 GB, and supports flexible inclusion or exclusion of sources, allowing the user to scope the graph to a domain of interest. Built-in analytical tools include multi-hop reasoning, subgraph extraction, text, vector and hybrid entity search, and enrichment analyses, all accessible through an interactive web interface, a REST API, and a Model Context Protocol server, the last enabling direct consumption by agentic and LLM-based systems. KRAKEN is freely available at https://app.krakenkg.com.
Zhou, Y.; Huang, F.; Zhao, Y.
Show abstract
Tuberculosis remains a major global public health threat. While whole-genome sequencing has transformed our understanding of the causative agent, Mycobacterium tuberculosis (MTB), existing genomic databases are highly fragmented and often underrepresent structural variations (SVs). Furthermore, critical population-genetic statistics are rarely integrated with phylogenetic and geographic context, forcing researchers to reconcile separate datasets manually. To address this gap, we developed TBpop (https://tbpop.chinacdc.cn), an open-access, integrated population genomics portal. TBpop is built from 420 clinical MTB isolates selected from the first national drug resistance baseline survey in China. The portal integrates isolate metadata, pangenome categories, SNPs, SVs, IS6110 insertion sites, strain phylogeny, and gene-level statistics, and provides three interactive explorer modules: the Population Explorer, the Statistics Explorer, and the Variation Explorer. Additionally, a User Analysis module allows researchers to run population genetic workflows on their own alignments. TBpop provides an integrated platform for exploring genome plasticity, signatures of positive selection, and conservation patterns of functionally important genes in MTB.
Murad, A. B.
Show abstract
BackgroundReuse of archived transcriptomic data underpins a large and growing share of published genomics. Because differences in upstream processing confound cross-study comparison, uniform reprocessing compendia -- recount3, ARCHS4, DEE2, refine.bio, Expression Atlas -- are widely treated as the remedy, and their availability is routinely assumed at the point of study design. Whether that remedy is actually obtainable for the population of published disease RNA-seq has not been measured. Prior audits have characterised metadata completeness and deposition rates, but none has quantified, across the published population, what fraction of studies can be uniformly reprocessed or where in the path from publication to comparable counts that capability is lost. ResultsWe enumerated 1,124 MeSH disease descriptors exhaustively, retrieved 16,820 human RNA-seq series from the Gene Expression Omnibus, and audited the 3,631 bulk, Illumina-platform series of at least 25 samples under three independently pre-registered, tool-enforced analysis plans. Raw reads were publicly available for 94.1% of series under a dual-route evidence standard, but only 46.7% appeared in any uniform reprocessing compendium (bounds 46.7-57.2%) and only 27.3% were usably covered at a 90% run threshold (bounds 27.3-38.8%). Of the 3,418 series whose reads are public, 991 were usably covered, leaving 71.0% of read-public series reprocessed by nothing usable. Presence overstated usability: DEE2 was present for 30.2% of series but usable for 3.4%. Design attrition was independent and severe -- 24.7% met bulk primary-tissue case-control criteria, 7.4% additionally reached a minimum replication threshold counted on sample accessions, and 4.1% did so counted on distinct donors. Among the 199 series where donor identity resolves, 4.1% pass the replication criterion on donors against 15.1% on accessions, a 3.73-fold difference; across the census frame, accessions exceeded distinct donors by 2.65-fold (Manski bounds 1.09-10.77, Imbens-Manski 95% CI 1.07-11.88). Independently, 36.4% of series-to-disease attributions produced by a conventional keyword query were refuted by the curated MeSH headings of the series own linked publication. ConclusionsUniform reanalysis of published human disease RNA-seq is unavailable for most studies in the population audited, and the binding constraint is usable coverage rather than deposition of raw reads. The loss occurs at several independent layers with different remedies, and the coverage layer -- unlike the others -- is one that resource maintainers can act on. Automated retrieval further over-counts eligible studies, both by admitting designs outside scope and by assigning studies to diseases their publications do not support.
Tindall, C.; Long, R. A.; Naughton, B.; Mapes, B. M.; Vismer, D.; Skinner, H. G.; Malenfant, J.; Maurya, M. R.; Nalls, M. A.; Ramachandran, S.; Nguyen, T.; Peters, M. A.; Scheuermann, R. H.
Show abstract
SysBio FAIRplex is a Common Fund Venture Program that catalogs and indexes data from the Accelerating Medicines Partnership(R) (AMP(R)) Program through a federated model in which data hosts retain custody of their datasets. The central piece of this work is the SysBio Common Data Model (SysBio CDM). AMP is a precompetitive public-private partnership started in 2014 that unites the resources of NIH and private partners to improve our understanding of disease pathways and transform current models for developing new treatments by: - identifying new targets, biomarkers, and development paradigms; - developing leading-edge tools and technologies; - collecting large-scale datasets and supporting analytics for open analysis by the public; and - generating consensus platforms and procedures. A multidisciplinary Task Force was chartered to design the SysBio CDM by extending the Observational Medical Outcomes Partnership (OMOP) Common Data Model into the -omics domain. The Task Force produced a Minimum Viable Product comprising nine OMOP tables; four extension tables for assay and file metadata; and a Common Data Element (CDE) Registry to specify field semantics. This manuscript describes the deliverable: the underlying design choices, the criteria applied in selecting and constructing the extension tables, how the extended model supports multimodal data integration across AMP projects, and what further work to support additional -omics modalities would entail. As an auxiliary methodology, the paper also describes the AI-assisted CDE harmonization workflow used to populate the model.
Medina-Ortiz, D.; Olivera-Nappa, A.; Lienqueo, M. E.; Opazo, R.; Romero, J.
Show abstract
Bacteriophage lytic enzymes and depolymerases are relevant to phage biology, antimicrobial development, and protein engineering, but their sequence and annotation data remain dispersed across general databases, specialized resources, genome-centred collections, and prediction-oriented datasets. We present PhageLysData, an evidence-aware and AI-ready resource constructed through reproducible multisource integration, provenance tracking, and exact-sequence consolidation. The release integrates 807,366 source observations from seven primary resources into 759,105 unique exact-sequence entities, comprising an evidence-supported Core of 11,867 entities, a Prediction Extension of 745,092 prediction-only candidates, and 2,146 Context entities retained for provenance and reference. This architecture preserves broad sequence-space coverage while maintaining a clear distinction between non-predictive and prediction-derived support. Core entities are enriched with harmonized biological annotations, physicochemical properties, independent InterProScan-derived functional annotations, mapped PDB and AlphaFold DB structural assets, and reusable numerical representations. For 11,259 eligible Core sequences, PhageLysData provides embeddings from 11 protein language models together with one-hot encoding under a common representation contract. Release-facing examples demonstrate latent-space exploration, unsupervised clustering, supervised classification, and evidence-aware candidate retrieval without defining a universal predictive benchmark. PhageLysData provides a traceable, versioned, and computationally accessible foundation for protein retrieval, comparative analysis, task-specific dataset construction, and machine-learning applications involving phage lytic enzymes and depolymerases.
Chi, L. A.; Ytreberg, F. M.
Show abstract
Protein-peptide interactions are central to cellular signaling and to a growing class of peptide therapeutics, yet the datasets used to develop and benchmark computational methods for protein-peptide modeling remain poorly standardized. Available databases prioritize comprehensive coverage but require task-specific curation, while published benchmarks are typically distributed as static collections built with heterogeneous curation, quality-filtering, redundancy-reduction, and sampling strategies, limiting reproducibility and cross-study comparison. We present PepXPro, a modular framework that transforms publicly available protein-peptide structure-affinity resources into curated datasets and reproducible benchmark collections generated under user-defined criteria. PepXPro is organized into three components: Scrape, for deterministic curation of protein-peptide complex entries from public resources; GenSample, for constructing configurable subsets under explicit quality, redundancy, and sampling constraints; and Benchmark, for evaluating candidate subsets and selecting a nonredundant, representative, general-purpose benchmark for distribution. Starting from PDBbind and complementary resources, the curation pipeline yields a pool of proteinpeptide complex entries that retains chemically complex cases, including disulfide- linked cyclic peptides, which are commonly excluded from existing benchmarks. We release PepXPro Benchmark v1, a benchmark comprising 70 non-redundant protein- peptide complexes with experimentally determined structures and binding affinities. The underlying framework provides an extensible foundation for reproducible protein- peptide benchmark construction.
Xuan, H.; Pasupuleti, R.; Liu, B.; Sun, H.; Zhang, J.; Yao, Z.; Zhong, C.
Show abstract
Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale. The lack of a comprehensive and up-to-date catalog of bioinformatics resources hinders efforts toward automated biomedical knowledge extraction and streamlined data analysis. Here we present SNAIL, a hybrid named entity recognition framework designed to automatically identify bioinformatics software and database (SW/DB) names from biomedical texts. SNAIL integrates complementary lexical and semantic modeling strategies. The lexical component captures orthographic patterns and contextual cues characteristic of SW/DB names, while the semantic component leverages contextual embeddings generated by transformer-based language models such as SciBERT, combined with an explicit token-masking strategy to enhance entity-focused representations. A large training corpus was constructed automatically through a hybrid pipeline that integrates citation-hinted extraction with large language model-assisted distillation. Evaluation on two independent benchmark datasets and real-world research articles demonstrates that SNAIL substantially outperforms existing approaches, including domain-specific methods such as bioNerDS2 and general-purpose large language models such as ChatGPT, Gemini, Grok and Claude. Applying SNAIL to large-scale literature analysis further reveals distinct journal-level preferences across bioinformatics subfields. These results demonstrate that SNAIL provides an accurate and scalable solution for identifying bioinformatics resources in scientific texts and enables systematic meta-analysis of tool usage and research trends.
McLaughlin, J.; Puig-Barbe, A.; Ibrahim, A.; Pava, D.; Pendlington, Z. M.; Matentzoglu, N.; Sollis, E.; Foreman, A.; Wilson, R.; Lopez Gomez, F.; Harris, L.; Adeleye, Y.; Kaur, S.; Meldal, B.; Smedley, D.; Parkinson, H.
Show abstract
Researchers increasingly need to explore hypotheses that span multimodal data across different scales, organisms, and domains. In practice, this requires connecting knowledge across fragmented databases with incompatible APIs and heterogeneous annotation practices. Large language model (LLM) agents can automate this data integration process, but grounding LLM agent outputs in scientifically correct sources of truth remains a significant challenge. Here we describe our deployment of a novel AI semantics workflow using LLM agents to enable scalable data integration, grounded in biological knowledge in the form of ontologies. Our workflow comprises (1) a multi-agent system curating scientific knowledge across ontologies using the Ontology Lookup Service (OLS) as grounding; (2) an LLM embedding service to enable interoperability between scientific databases by mapping ontology terms; and (3) GrEBI, a knowledge graph and Model Context Protocol (MCP) server enabling LLM agents to conduct cross-cutting, multi-omic biomedical queries. O_FIG O_LINKSMALLFIG WIDTH=187 HEIGHT=200 SRC="FIGDIR/small/742514v1_ufig1.gif" ALT="Figure 1"> View larger version (39K): org.highwire.dtl.DTLVardef@1fc525aorg.highwire.dtl.DTLVardef@82c4f7org.highwire.dtl.DTLVardef@15173fdorg.highwire.dtl.DTLVardef@961f1c_HPS_FORMAT_FIGEXP M_FIG C_FIG
Baker, M.; Bett, K.; Vargas, A.; Jin, L.
Show abstract
Structural variants (SVs) are large-scale genomic variants, which can disrupt important functional and regulatory elements, leading to genomic disorders in humans and playing important roles in domestication, disease resistance, and traits in plants. SVs are generated across populations of individuals and used for association studies, consisting of large datasets with thousands of genomic loci. Visualization of these SVs aids in understanding their genomic distribution, identifying patterns across affected or phenotypic groups, and assessing their proximity to other genomic regions of interest. A variety of tools exist for visualizing SVs, including linear genome browsers and graph-based methods; however, many do not offer intuitive or scalable representations of SVs across large populations. To address this, we present SVPopEx, an interactive tool for population-wide visualization and exploration of SVs. SVPopEx provides a unique and intuitive representation for insertions, deletions, inversions, duplications, and translocations in a linear genome-style browser. Novel features were developed to support comparisons across genomes within user-defined regions, including rendering SVs based on one or more samples and visualizing haplotypes. Use of the tool is demonstrated with SV datasets from Schistosoma mansoni and Lens culinaris. A task-based evaluation was conducted using SVPopEx and two other linear genome browsers, which demonstrated that SVPopEx excelled in (1) providing a clear representation of the SVs present and (2) supporting comparisons across genomes.
Pozhidaeva, M.; Schreiber, S.; Schubert, K.; Busch, W.; Hackermüller, J.; Canzler, S.
Show abstract
Toxicological omics studies require comprehensive metadata to support reproducibility, interoperability, and regulatory reuse. However, metadata requirements differ across public repositories, reporting frameworks, and laboratory workflows, resulting in inconsistent annotation and limited data integration. To address this challenge, we developed MeSyTo (Metadata for Systems Toxicology), an ontology-driven framework for harmonizing metadata across toxicological omics. Metadata concepts from public repositories, the OECD Omics Reporting Framework (OORF), community standards, and institutional workflows were semantically aligned and implemented as the MeSyTo Metadata Model (MMM). The MMM serves as the basis for the automatic generation of SHACL validation shapes and framework-specific metadata profiles, while curated value sets are represented as SKOS controlled vocabularies to support metadata collection and validation. The current implementation comprises 105 ontology classes and 527 data properties and supports transcriptomics, proteomics, and metabolomics. A prototype web application demonstrates ontology-driven metadata collection with integrated semantic validation and ontology-based term resolution. The ontology, validation shapes, controlled vocabularies, generation scripts, and software are publicly available as open-source resources. MeSyTo provides a reusable semantic foundation for harmonized, machine-actionable metadata and facilitates repository submission, regulatory reporting, and interoperable data exchange across toxicological omics studies.
Gupta, S.; Verma, A. K.; Jana, S.; Ahmad, S.
Show abstract
Abstract Background: Legacy microarray datasets provide an extensive record of human transcriptomic biology, but their reuse is constrained by differences in platform design, preprocessing, measurement scale, and gene coverage. Platforms measuring only subsets of genes cannot readily be integrated with higher-coverage platforms, limiting large-scale analysis and computational modeling. Results: We developed Multi-Platform Gene Expression Matrix (MPGEM), a computational framework and resource for harmonizing and completing gene-expression profiles across heterogeneous microarray platforms. MPGEM uses a Reference Quantile Distribution (RQD) and generalized Reference Subset Quantile Distribution (RSQD) framework to transform profiles with different gene coverage onto a common quantitative scale. The MPGEM Engine, a multilayer perceptron, predicts expression of unmeasured genes from genes shared across platforms. Applied to Affymetrix GPL570, GPL571, and GPL96, MPGEM uses GPL570 as a 19,320- gene reference space comprising 12,712 predictor and 6,608 target genes. The resulting resource contains 207,135 human gene-expression profiles across 19,320 genes. Evaluation using masked GPL570 profiles yielded mean sample-wise Pearson and Spearman correlations of 0.944 and 0.939, respectively, and mean gene-wise correlations of 0.830 and 0.825. The lowest-performing 5% of target genes achieved a mean Pearson correlation of 0.683. MPGEM showed comparable or higher predictive performance than baseline mean imputation and K-nearest-neighbor approaches. Conclusions: MPGEM transforms heterogeneous, partially measured legacy microarray profiles into a harmonized, transcriptome-complete representation, facilitating their reuse for large-scale transcriptomic analysis, biomarker discovery, systems biology, and machine learning. The framework, trained models, and expression resource are provided as open-source resources.
Nguyen, T. Q.; Do, K. H. D.; Vu, T. M.; Hoang, N. V.
Show abstract
Khang Dan 18 (KD18) is an Oryza sativa L. subsp. indica rice cultivar widely cultivated in northern Vietnam and used as an experimental and breeding background in Vietnamese rice research. Although KD18 has previously been represented in low-depth population resequencing datasets, a contiguous and annotated cultivar-specific genome has not been available. Here, we report a chromosome-scale genome assembly of KD18 generated using Oxford Nanopore long-read and Illumina short-read sequencing. The 395.3-Mb assembly comprises 12 chromosome-scale pseudomolecules containing approximately 95% of the assembled sequence and 99.6% of the predicted protein-coding genes. The assembly showed 97.2% BUSCO completeness, an average Merqury quality value of 46 and a long terminal repeat assembly index of 13.21. A total of 56,546 protein-coding genes representing 71,237 transcripts were predicted, with 99% BUSCO and 98.68% OMArk completeness. These statistics are similar to those of other high-quality genome assemblies that were recently published for different Asian rice cultivars, therefore providing a cultivar-specific genomic resource for research involving KD18 and KD18-derived materials.
Barbosa Araujo, P. V.; da Silva Fiuza, T.; Ferraz, R. S.; Kroll, J. E.; Andrade, R. L.; Gomes, D. H. F.; Varuzza, L.; de Souza, G. A.; de Souza, S. J.
Show abstract
Polygenic risk scores (PRS) have emerged as a powerful tool for quantifying genetic susceptibility to complex traits and diseases. However, their calculation and interpretation require standardized data curation, robust statistical methods, and clear reporting strategies. In this work, we present an integrated pipeline designed to address these challenges. The pipeline begins with the construction of a curated genotype/phenotype database derived from public repositories, ensuring that only phenotypes with appropriate metadata, statistical distributions, and ethical suitability are retained. The final dataset comprises 2,346 phenotypes covering 38,256,468 unique SNPs. These phenotypes serve as the final analytical units for PRS calculation, risk stratification, and individual-level interpretation. The generated reports integrate sample-level results, phenotype categorization, risk classification, study references, and variant tables, providing a structured and interpretable output for end users. Together, the curated database and reporting framework establish a comprehensive toolbox for PRS analysis, enhancing reproducibility, transparency, and usability in both research and clinical contexts.
Mainguy, J.; Lemane, T.; Bazin, A.; Arnoux, J.; Gautreau, G.; Medigue, C.; Calteau, A.; Vallenet, D.
Show abstract
PanGBank (https://pangbank.genoscope.cns.fr) is a comprehensive open-access database providing precomputed prokaryotic pangenomes at a broad taxonomic scale. Built upon PPanGGOLiN partitioned pangenome graphs, PanGBank addresses the growing need for large-scale comparative genomics through a standardized, regularly updated, and fully accessible resource. The initial release comprises two complementary collections covering more than 4,600 prokaryotic species from the Genome Taxonomy Database (GTDB), encompassing over 393,000 genomes: GTDB all, maximizing taxonomic and environmental diversity through the inclusion of MAGs and SAGs, and GTDB refseq, focusing on high-quality, annotation-rich genomes. Each species-level pangenome integrates graph-based statistical partitions into persistent, shell, and cloud gene families, together with regions of genomic plasticity (panRGP) and co-localized functional modules (panModule). PanGBank offers multiple access modes, including a REST API, a command-line interface (PanGBank-cli), and an interactive web interface. By combining large-scale pangenome resources with advanced graph-based analyses, PanGBank provides a scalable framework for exploring microbial diversity, genome evolution, functional variation, and the dissemination of adaptive traits across prokaryotic populations, as illustrated by a use case on Acinetobacter baumannii pangenome investigating the distribution and evolution of antimicrobial resistance determinants. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=63 SRC="FIGDIR/small/742796v1_ufig1.gif" ALT="Figure 1"> View larger version (23K): org.highwire.dtl.DTLVardef@186b88dorg.highwire.dtl.DTLVardef@1be33d0org.highwire.dtl.DTLVardef@3bc596org.highwire.dtl.DTLVardef@292401_HPS_FORMAT_FIGEXP M_FIG C_FIG
Aires Teixeira, J. V.; Motta Venancio, T.; Quintanilha-Peixoto, G.; Pimenta de Oliveira, K. K.
Show abstract
MicroRNAs (miRNAs) are key post-transcriptional regulators of development, stress response, and secondary cell wall formation in woody plants, yet annotations for Eucalyptus grandis, the world's most widely planted hardwood, remain fragmented across studies using incompatible discovery pipelines and filtering criteria. Here we present the Eucalyptus MicroRNA Archive (EMA), a curated, locus-resolved database integrating three independent small RNA sequencing datasets spanning vegetative tissue, somatic embryogenesis, and mechanically induced tension wood formation. Applying annotation criteria aligned with current plant miRNA standards, EMA catalogs 99 curated miRNAs (31 previously described, 68 novel) organized into 34 family-level groupings under a three-tier confidence system, known-reference-supported, multi-study replicated, or single-study, that preserves study-of-origin and sample-level evidence for every entry. Cross-study comparison showed that only 9 of 99 entries (9.1%) were independently supported by all three datasets, supporting an evidence-tiered rather than binary annotation scheme. Target prediction against the E. grandis transcriptome yielded 1,773 miRNA-target interactions spanning 764 loci, integrated into a combined miRNA-target and protein-protein interaction network. This network resolved into functionally coherent, mutually isolated clusters, including an miR482-associated NBS-LRR/TIR disease-resistance hub with a substantial translational-repression component, alongside modules enriched for ribosome biogenesis and translation, DNA replication, and nitrogen and carbohydrate metabolism. EMA is publicly accessible through an interactive web dashboard, with all curated data, source code, and analysis scripts openly available, providing a reproducible, extensible framework for E. grandis miRNA research and a template for similarly structured resources in other non-model woody species.
Qi, Y.; Lundy-Perez, K.; Gee, D. A.; Chambwe, N.
Show abstract
Objectives Accurate phenotyping of cases and controls is essential for studying biological and environmental contributors to disease in large biobanks. We aimed to develop a flexible, customizable, and reproducible electronic health record (EHR)-based phenotyping framework for identifying disease cases and generating matched control cohorts for downstream analyses. Here, we developed the Phenotyping Algorithm for Cases and matched Controls using EHR-based Rules (PACER). Materials and Methods Applying PACER to the All of Us Research Program Curated Data Repository v8.0, we identified female breast cancer (BC) cases identified among participants recorded as female at birth using at least two BC-associated diagnostic Observational Medical Outcomes Partnership concept IDs documented at least 30 days apart. A one-to-one matched control cohort was generated by jointly matching on sex, age, genetic ancestry, and state-level residency. Clinical, socioeconomic, and genomic data were integrated for analysis. Results We identified 10,225 BC cases and generated a control cohort of the same size matched for key demographic characteristics. Comparison with a phecodeX-based BC cohort showed 91.03% agreement. Among cases responding to relevant survey items, 80.86% self-reported a personal history of BC, compared to 1.89% of controls. We detected an enrichment of BC-associated GWAS catalog variants, pathogenic mutations in known risk genes, and higher polygenic risk scores in cases compared to controls. Discussion and Conclusion Concordance across a phecodeX-based cohort, self-reported survey responses, and genomic analyses supports the validity of PACER-defined cohorts. PACER is publicly available and readily adaptable to other diseases, supporting future research in risk modeling and precision medicine.
Kaniewski, P.; Carter, E. K.; Rhodes, D.; Lim, E. M.; Li, J.; Vergine, J.; Matentzoglu, N.; Schaper, K.; Reilly, J.; Sundar, S.; Vijnck, L.; Sharp, E.; Alfonso, N.; Ford, A.; Stepanenko, A.; Hempstead, C.; Brokmeier, P.; Bizon, C.; Tropsha, A.; Haendel, M. A.; Fajgenbaum, D. C.; Lancashire, L.
Show abstract
Identifying causal connections between existing drugs and mechanistic profiles of diseases is a foundational step for effective drug repurposing. Although knowledge graphs (KGs) are highly suited for consolidating biomedical databases and tracking these connections, a single biomedical KG is constrained by its ingestion pipeline and knowledge sources. While different biomedical KGs could be complementary if combined, efforts to combine them into a unified and more comprehensive KG are hindered by lack of interoperability and poor provenance. To address those issues, we present EC-KG, a Biolink Model-compatible KG for computational drug repurposing. EC-KG is an interoperable, provenance-first KG which integrates RTX-KG2, ROBOKOP, and PrimeKG at the network-level, encapsulating over 7 million nodes and 81 million edges from 95 primary data sources. EC-KG has improved coverage of core biomedical entities such as drugs, targets, and diseases relevant to drug repurposing vs source graphs, and captures complex biomedical mechanisms within its topology. We demonstrate that the network unification in EC-KG leads to emergence of novel, mechanistically relevant pathways which are disconnected in the underlying constituent networks and show its applications in method development, benchmarking and predictive drug repurposing applications. EC-KG has already been successfully used in drug repurposing research to surface Botulinum Toxin A as a candidate to treat Major Depressive Disorder, as well as to validate repurposing of Lenalidomide and Dexamethasone for a subgroup of patients with Rosai-Dorfman Disease.
Motta, J. A.; Motta, M. d. M.; Fernandez, C.
Show abstract
In this work, we present a machine learning model for identifying pathogenic DNA variants. The model was learned from the analysis of normal and pathogenic sequences extracted from the ClinVar database (supported by NCBI). This analysis was based on a conceptual semantic model of DNA sequences converted to peptide sequences (amino acid sequences) governed by a well-defined grammar, which allowed us to apply NLP techniques, specifically Part of Speech tagging (POS tagging). Our predictive model was built by combining two techniques: CRF (from the Markov model family), which performs the sequencing, and BiLSTM (a deep learning model) which captures the past and future content of the sequences. The training space was created with the sequences of 105 genes associated with approximately 27,000 pathogenic variants. The model was evaluated using the metrics precision, P-R and ROC curves, AUC, and confusion matrices. Its performance was also compared against five known methods for predicting pathogenic variants. The results show exceptional performance that exceeds expectations and places this new method at the state of the art for predicting pathogenic DNA sequences.
Arif, A.; Filho, J. V. d. S.
Show abstract
The increasing use of tumor sequencing has intensified the need for fast, traceable interpretation of genomic variants. General-purpose large language models can produce fluent answers, but unsupported statements, weak provenance, and stale knowledge limit their suitability for clinical genomics. We developed OncoGenRAG, a research framework that combines a parameter-efficiently fine-tuned BioBERT classifier with an entity-aware retrieval system over a curated, multi-source oncology knowledge base. The reported knowledge base contains 933 harmonized records derived from CIViC, ClinVar/dbSNP, Open Targets, UniProtKB/Swiss-Prot, Ensembl Variation, and linked PubMed literature. The classifier assigns one of five labels: Pathogenic, Likely Pathogenic, Variant of Uncertain Significance, Benign, or Oncogenic; the retrieval component ranks evidence records using subword TF-IDF similarity and explicit gene, variant, and cancer-type matches. A rejection rule suppresses answers when retrieval support is below a prespecified threshold. In the authors held-out evaluation, the classifier achieved 92.40% accuracy, 93.15% weighted precision, 92.40% weighted recall, and 92.65% weighted F1 score. In a separate benchmark of 100 clinical-style queries, OncoGenRAG achieved reported Precision@1 of 94.5%, Precision@3 of 96.8%, and 100% database grounding. No hallucinated answer was observed under the study operational definition, compared with a 41.0% no-hallucination rate for the ungrounded baseline. These results should be interpreted as internal validation rather than proof of universal safety because query construction, annotator agreement, class-specific performance, calibration, and external validation data were not available for independent analysis. OncoGenRAG provides a transparent design for evidence retrieval and abstention, but it is a research prototype and must not be used to select treatment without expert review.