Bioinformatics
◐ Oxford University Press (OUP)
Preprints posted in the last 90 days, ranked by how well they match Bioinformatics's content profile, based on 1204 papers previously published here. The average preprint has a 0.85% match score for this journal, so anything above that is already an above-average fit.
Ahn, S.; Oh, E. J.; Prada, D.; Shojaie, A.
Show abstract
Recent advances in spatial proteomics, particularly imaging mass cytometry, enable the measurement of protein expression at the single-cell level while preserving a spatial context. Conventional survival analyses, however, typically rely on patient-level averages of protein intensities and therefore overlook spatial heterogeneity and tissue architecture. To address this limitation, we introduce a framework that incorporates spatial information into survival modeling by generating spatially adjusted protein summaries (SAPS). In this approach, cell-level protein intensities within each patient are modeled using spatial spline regression to capture spatial trends. From these models, we extract two complementary features: a spatially adjusted mean expression and a residual variance that reflects cell-to-cell variability unexplained by spatial effects. These summaries are then incorporated into Cox proportional hazards models in combination with clinical covariates. In simulation studies, our proposed framework achieved improved predictive performance compared to other alternative methods. The application of the method to breast cancer imaging mass cytometry data indicate that spatially adjusted summaries may enhance survival prediction and reveal biologically interpretable spatial protein patterns, suggesting high translational potential. This methodology offers an efficient means of translating complex spatial proteomics data into patient-level features, providing both improved survival prediction and new insights into the role of spatial heterogeneity in cancer outcomes.
Ashrafiyan, S.; Kosaretskii, I.; Schulz, M. H.
Show abstract
MotivationThe rapid expansion of single-cell RNA sequencing (scRNA-seq) atlases has generated datasets comprising millions of cells annotated with increasingly rich metadata, including tissue, cell type, disease status, sex, age, treatment, and temporal information. Biological questions frequently require simultaneous interrogation of multiple metadata dimensions, such as identifying specific cell populations within defined tissues, disease states, demographic groups, and time points. While existing interactive platforms facilitate visualization and analysis of scRNA-seq data, deep metadata-driven exploration and downstream analysis of atlas-scale datasets remain insufficiently supported. ResultsWe developed AtlasLens, an open-source R/Shiny application for interactive exploration of scRNA-seq datasets and integrated cellular atlases. AtlasLens enables iterative filtering across arbitrary metadata combinations, allowing users to define biologically meaningful cellular subsets and immediately perform downstream analyses. The platform integrates interactive visualization, differential expression analysis, Gene Ontology enrichment with redundancy reduction, temporal expression analysis, and context-dependent gene function profiling through GeneCOCOA. AtlasLens additionally records analysis history and automatically generates corresponding R code to enhance reproducibility. The application is distributed through Docker for simple local deployment, preserving data privacy and eliminating dependency-management challenges. We demonstrate AtlasLens using the Tabula Muris and a time-resolved whole-lung single-cell atlas of bleomycin-induced lung injury and fibrosis, highlighting its ability to support complex metadata-driven biological investigations. AvailabilitySource code is available at https://github.com/SchulzLab/AtlasLens. Contactmarcel.schulz@em.uni-frankfurt.de
zhang, X.
Show abstract
MotivationConstraint-based metabolic modeling faces a calibration gap: genome-scale metabolic models (GEMs) integrated with transcriptomics alone rely on expression-to-flux heuristics (E-Flux, GIMME, iMAT, MOMENT) that ignore cross-omics co-variance structure and lack statistical mechanisms for propagating omics uncertainty into reaction bounds, yielding flux predictions with limited agreement to 13C metabolic flux analysis (MFA) measurements. ResultsWe present Chemo-Calib, a multiblock PLS (MB-PLS) framework that calibrates GEM reaction bounds from the shared latent structure of metabolomics, transcriptomics, and proteomics data. On 11 E. coli 13C-MFA reference conditions spanning the Keio fluxome and Holm 2010 datasets, ChemoCalib constrained FBA on iJO1366 achieves a Spearman{rho} = 0.461 overall (up to 0.523 in PPP) and Pearson r of 0.49-0.58 across central carbon pathways, with statistically significant improvement over expression-only baselines including E-Flux2 and SPOT (p < 0.05, Holm-corrected). The latent-to-constraint mapping employs GPR-aware VIP aggregation (Algorithm 1) to project multi-omics latent scores onto genome-scale reaction bounds without heuristic thresholding. An optional in-silico active learning loop (relegated to Supplementary Material) further tightens calibration through virtual experiment selection. AvailabilityChemoCalib is open-source (MIT) at https://github.com/chemocalib/chemocalib with Docker support, a 5-minute tutorial, and pre-computed iJO1366 benchmarks. Preprint available at bioRxiv; code archived at Zenodo DOI: 10.5281/zenodo.21645890.
Shah, R. N.; Ouyang-Zhang, J.; Cohen, Z.; Briglia, M. R.; Zhang, C.; Klivans, A.; Diaz, D. J.
Show abstract
Accurate modeling of antibody-antigen (Ab-Ag) complexes is central to biologic development, yet the reliability and failures of modern Ab-Ag folding pipelines remain poorly characterized. Single-chain variable fragments (scFvs) are therapeutically important antibodies, but large-scale evaluations of structure prediction models on scFv-Ag complexes are largely lacking. We introduce a scalable benchmarking pipeline that generates large ensembles of scFv-Ag structure predictions by cofolding a curated subset of 3,800 Ab-Ag complexes from SAbDab using multiple state-of-the-art models under diverse inference-time settings. The resulting dataset, SCALE (scFv-Ag CompLex Ensembles) includes standardized scFv-Ag sequences and around 200,000 predicted complexes spanning different models, sampling strategies, and auxiliary inputs. Using SCALE, we evaluate model performance in recovering correct scFv-Ag interfaces and assess the ability of existing confidence metrics to select the best structure from prediction ensembles. We find that while confidence scores effectively distinguish easy from hard scFv-Ag complexes, they often fail to identify the highest-quality interface for a given target. Further analysis shows that near-correct interfaces typically appear in ensembles but at low frequency, and inference-time choices like sampling, recycling, and using evolutionary or structural information are crucial for accurate scFv-Ag complex predictions. Dataset and analysis code are available at https://huggingface.co/datasets/ravishah1/SCALE
yang, c.; Cook, N.; Zeng, Y.; Fu, T.; budde, J.; Cruchaga, C.; Belloy, M. E.
Show abstract
Summary It has become standard practice to visualize regional signals from genomewide association studies GWAS using LocusZoom plots Similarly GWAS signals are compared to regionally matched quantitative trait loci QTLs ie varianttogene regulation data using LocusCompare plots to aid assessment of candidate traitrelated genes Despite broad usage these tools annotate variants by linkage disequilibrium LD to a single lead or index variant This singleindex representation has limitations for visualizing complex loci that contain multiple independent signals We present LocusBlend an interactive web application for multiindex LDblended visualization of genomic loci LocusBlend supports one or two genomic association summarystatistic datasets and one to three index variants multiindex LocusZoom colorblended plots and matching LocusCompare visualizations Applications to Alzheimers disease GWAS and QTL signals illustrate LocusBlend enables visualization and separation of independent signals despite shared LD and high genomic complexity Overall LocusBlend is aimed at supporting researchers handle the continuously expanding complexity of human genomics findings Availability and Implementation LocusBlend is freely available at httpslocusblendwustledu Publication ready plots are generated in 1min Source code documentation example datasets input templates and reproducibility instructions are available at httpsgithubcomBelloyLabLocusBlend LocusBlend is implemented in Python using Streamlit Plotly and PLINK Supplementary Information Supplementary data are available online
Jia, P.; Chu, L.; Ren, Z.; Cui, H.; Shao, B.
Show abstract
MotivationHigh-resolution spatial transcriptomics enables single-cell ligand-receptor analysis, but unmasked receptor spatial lags include receptor expression from non-target neighbors, complicating the attribution of local communication signals to specified source-target cell-type pairs. ResultsWe present MaskTalk, a Python package implementing the cell-identity-gated spatial lag model (CIG-SLM). CIG-SLM restricts receptor-side neighborhoods to target cells through W(C) = WD(C). In cell-level breast cancer Visium HD data, CIG-SLM produced target-cell-dependent communication profiles relative to the matched LIANA+ bivariate unmasked baseline, and masked-specific records showed larger between-condition effect sizes and shorter physical source-target distances. Public breast cancer Xenium data demonstrated that MaskTalk runs on external cell-level spatial data and exhibits target-aware masking behavior. Availability and ImplementationImplemented in Python with AnnData; code is available at https://github.com/JiaPP1994/MaskTalk. Software and data archive DOIs are 10.5281/zenodo.21735883 and 10.5281/zenodo.21735951, respectively. Contactshaobin79@aliyun.com; cuihw2001423@163.com Supplementary InformationSupplementary Methods S1-S4, Figures S1-S3, and Tables S1-S5.
Douglas, J. M.; Lynch, B. J.; Yiu, J. C. H.; Nicholson, S.; Vasquez-Rios, C.; Ma, D.; Huntsman, D. G.; Park, Y.
Show abstract
SummaryA modular FASTQ-to-figures solution for analyzing low-depth or shallow whole-genome sequencing (sWGS) data. Shallow WGS can be used to detect copy number (CN) aberrations, Homologous Recombination Deficiency (HRD), and to detect and create CN signatures. Growing in popularity, this sequencing type is used for neonatal diagnostics and studying cancer. One of the major benefits is the reduced cost compared with deeper sequencing modalities such as Whole Genome Sequencing (WGS). With just 15 million reads often targeted, and sample substrate options like Formalin-Fixed, Paraffin-Embedded (FFPE) blocks widely available, this approach enables an affordable study to be performed at scale. Our pipeline and R package are an end-to-end solution implemented with reusability and modularity in mind. It makes entry and exit from the ecosystem easy, providing regular standardized output formats throughout execution. The pipeline is written in the well-supported and cross-platform Nextflow framework and has been submitted for inclusion in nf-core. Additionally, a Docker image for the utanos R package has been created to improve modularity. Availability and ImplementationThe latest version of all software is freely available on GitHub. For the full processing pipeline, visit: https://github.com/Huntsmanlab/swgs-processing-pipeline. For just the utanos R package, visit: https://github.com/Huntsmanlab/utanos.
Rodriguez, Z. B.; Guare, L.; Caruth, L.; Cardone, K. M.; Carson, C. C.; Cherlin, T.; Mohammed, S.; Gupta, H.; Kumar, R.; Keat, K.; Verma, S. S.; Verma, A.
Show abstract
SummaryElectronic health record (EHR)-linked biobanks generate unprecedented genomic and phenotypic datasets, but their scientific utility is constrained by data fragmentation across institutional silos and incompatible computing infrastructures, forcing researchers to rewrite ad-hoc scripts for each new environment. We present the PMBB Geno-Pheno Toolkit, a suite of modular Nextflow pipelines for biobank-scale association analyses. This note focuses on the toolkits SAIGE family of pipelines -- supporting genome-wide (GWAS), exome-wide (ExWAS), and phenome-wide (PheWAS) association testing -- together with the companion GWAMA and ExWAS meta-analysis pipelines that enable cross-biobank replication. All components are containerized (Docker/Apptainer) and orchestrated with Nextflow, allowing the same workflows to run unmodified on local HPC clusters, cloud platforms, and the All of Us Research Workbench. Complementary toolkit pipelines for PLINK-based GWAS, polygenic scoring, LD-based clumping, and phenotype harmonization are also available and briefly noted. AvailabilityThe PMBB Geno-Pheno Toolkit is freely available at https://github.com/PMBB-Informatics-and-Genomics/pmbb-geno-pheno-toolkit under MIT open-source license.
Kalson, L.; Sexton-Oates, A.; Drevet, G.; Fernandez-Cuesta, L.; Foll, M.; Alcala, N.
Show abstract
Motivation: Integrated multi-omic analyses have transformed our understanding of cancer biology, giving rise to data-driven molecular classifications that capture disease heterogeneity beyond conventional histopathology. Among these approaches, multi-omic factor analysis (MOFA), a multimodal extension of principal component analysis, has been widely used to identify sources of molecular variation across omic layers and classify samples into molecular groups. However, classifying query samples according to an existing MOFA-based classification remains challenging, as there is no validated computational method for projecting samples into pretrained MOFA latent factor spaces. Results: We present PACMOS, an R package that provides a generalizable approach to project query samples into pretrained MOFA latent factor spaces. We validate PACMOS using two cancer datasets with published MOFA-based classifications - lung neuroendocrine neoplasms and pleural mesothelioma - showing that PACMOS preserves the existing MOFA latent factor space while allowing to classify query samples. Availability and implementation: PACMOS is an open-source R package available on the IARC bioinformatics GitHub organization (submitted to Bioconductor) at https://github.com/IARCbioinfo/PACMOS and DOI in Zenodo: https://doi.org/10.5281/zenodo.20933824, along with installation instructions and a vignette with an application. Supplementary information: Supplementary data are available in separate files.
Pocas, G.; Umar, M.; Davis, O.; Hemberg, M.; Lamurias, A.; Lakatos, A.; Asif, M.
Show abstract
MotivationSingle-cell RNA sequencing (scRNA-seq) has become an attractive tool for studying complex diseases, in which transient cell states affecting diverse cell populations characterise disease development and progression. However, due to data sparsity and disease heterogeneity analysis is often challenging. With recent advances in machine learning, two widely used approaches have emerged for learning cellular representations: large-scale foundation models and biological knowledge-guided methods. Despite their complementary strengths, there is currently no unified workflow for systematically comparing and integrating these approaches. ResultsHere, we present scRepresenter, an open-source workflow for computing, integrating, and validating cellular embeddings derived from foundation models and biological knowledge-guided methods in the context of complex diseases. It consists of two components: a command-line workflow that computes cellular embeddings and performs downstream analyses, and an interactive Shiny application for visualizing and comparing the computed embeddings. scRepresenter supports four categories of cellular representations: (1) expression-based, (2) knowledge-guided, (3) foundation model-derived, and (4) hybrid embeddings that combine foundation model-derived representations with knowledge-guided representations. This approach takes a cell-by-gene count matrix as input and outputs an integrated object containing the computed embeddings. Then, this object can be uploaded into our interactive Shiny application to compare different embeddings. AvailabilityThe workflow is available at https://github.com/GuilhermePocas/scRepresenter ContactAL291@cam.ac.uk; MA2129@cam.ac.uk
Gautam, P.; Mitra, P.
Show abstract
Prediction of B-cell epitopes can assist in reducing costly wet-lab screening in vaccine design, diagnostics, and antibody discovery. However, current predictors often suffer from noisy labels, weak generalization, and structure-dependent workflows. Here we present EO_SCPLOWPIC_SCPLOWESM-GA, an efficient sequenceonly pipeline for linear B-cell epitope prediction. Positive and negative peptide examples are collected from IEDB, which provides experimentally tested epitopes and distinguishes positive and negative epitope records based on assay evidence(Vita et al., 2019). Each peptide is encoded with a frozen ESM-2 protein language model: a bidirectional transformer producing amino acid embeddings for downstream structure and function tasks (Lin et al., 2023). Mean-pooled embeddings are further compressed into a compact 420-feature representation with a genetic algorithm and classified with lightweight Random Forest, XGBoost, or MLP heads. This avoids foundation-model fine-tuning, reduces the number of trainable parameters, improves interpretability, and enables low-resource deployment. On an IEDB-derived benchmark, EO_SCPLOWPIC_SCPLOWESM-GA attains 0.880{+/-} 0.004 AUC-ROC, 0.852{+/-} 0.005 PR-AUC, 82.0 {+/-} 0.6% accuracy, 0.79 {+/-} 0.01 F1, and 0.74{+/-} 0.01 MCC, outperforming dense ESM-2 features and baselines LBCE-XGB, EpitopeVec, and BepiPred-2.0 (mean{+/-} std over five independent random seeds). The framework shows how frozen protein foundation models can enable pandemic preparedness, peptide vaccine prioritization, diagnostic antigen screening, and equitable computational immunology.
Zhang, L.; Demarco, A. G.; Ghafari, K.; Devlin, B.; MacDonald, M. L.; Roeder, K.
Show abstract
MotivationKinases regulate a multitude of protein functions, and their dysregulation is pivotal for many human diseases. Direct measurement of kinase activity, however, is often challenging; therefore, inferring activity from the behavior of their substrates is a widely adopted strategy. Nonetheless, traditional methods typically oversimplify the underlying network, ignoring that any particular substrate can be phosphorylated by multiple kinases. ResultsWe present LIKA, a likelihood-based framework for inferring kinase activity from phosphoproteomic data. By modeling the many-to-many structure of kinase-substrate interactions, LIKA achieves high efficiency, even with limited data, while capturing network complexity. Simulation and cell line analyses confirm the robustness and accuracy of LIKA. Importantly, analysis of a phosphoproteomic dataset from schizophrenia and control subjects reveals novel dysregulated kinases. Availability and ImplementationThe implementation code and publicly available data are provided at: https://github.com/lujingz/LIKA.
Schmitz, M.; Rauschning, L.; Kawaguchi, K.; Sikic, M.
Show abstract
Accurate haplotype phasing is essential for high-quality genome assembly, yet de novo phasing of complex genomes without parental data remains challenging. We formulate haplotype phasing as a node clustering problem with overlapping clusters on augmented unitig graphs, where nodes represent contiguous, non-branching DNA sequence fragments and two edge types can encode sequence overlap or Hi-C proximity information. We introduce a contrastive learning framework with a custom objective function and train a graph-transformer-based model, termed grapHiC, to phase paternal, maternal, and homozygous unitig nodes. grapHiC; is the first machine-learning-based method to perform reference-free haplotype phasing and the first approach to directly phase raw unitig graphs without prior simplification. We show that grapHiCaccurately clusters nodes on human-genome-scale graphs and that its predictions can effectively guide phased de novo genome assembly, producing human assemblies with contiguity and phasing quality comparable to the state of the art when integrated with the DipGNNome assembler.
Rich, J.; Pachter, L.
Show abstract
Summary: fastQpick is a command-line tool and Python library for sampling FASTQ reads with replacement. Sampling with replacement turns a single FASTQ file into an arbitrary number of bootstrap replicates, which enables uncertainty quantification and statistical analysis at the level of raw reads. This process answers questions such as how much an abundance estimate would change if the library were resequenced, or whether a low-abundance call is robust to the particular reads that were sequenced. fastQpick works efficiently on large libraries by streaming files in two passes by default: first to count reads and create a hash-based counter, and then to write the sample. It generates a full-size bootstrap replicate of a 500-million-read library in under 30 minutes with 9.4 GB of peak memory, with a low-memory mode that reduces the peak to 1.4 GB. A single-pass mode draws samples in a single read through the file, using O(1) working memory and producing an output size that is exact in expectation but not fixed. In a real yeast RNA-seq experiment, bootstrap replicates generated by fastQpick recover the sampling uncertainty of transcript abundance estimates, matching the analytic multinomial standard errors to within a few percent. Availability and implementationfastQpick is open source and freely available under the MIT license on GitHub at https://github.com/pachterlab/fastQpick and on PyPI (pip install fastQpick).
Yeo, K.; Kim, D.; Sim, J.; Lee, J.
Show abstract
MotivationLigand binding-site similarity search is a crucial step in drug discovery that reduces the conformational search space for docking and other downstream tasks by comparing a target protein against experimentally identified binding sites. Existing methods rely on either direct structural alignment or lossy compression of structural information, producing a trade-off between scalability and precision. ResultsWe propose LEN-Seek, a ligand binding-site search method based on a graph neural network (GNN)-driven variational autoencoder (VAE) that encodes the 3D structural and physicochemical context of a binding site into a probabilistic latent space, enabling similarity search within a low-dimensional vector space. A binding site is modeled as a graph of amino acid residues, with node features adopted from the protein language model, Ankh, and edges encoded as SE(3)-invariant (roto-translational invariant) geometric relationships, thereby avoiding expensive data augmentation or SE(3)-equivariant models. Compared to ProBiS, the purely geometric graph-clique based method, LEN-Seek successfully retrieves a substantial portion of similar binding sites with a roughly 3,400-fold lower per-comparison cost, demonstrating its potential as a scalable approach to template-based ligand binding-site search in large-scale protein structure databases. Supplementary informationSupplementary data are available at Bioinformatics online.
Dong, Y.; Li, N.; Chiribau, C. B.; Mitchell, M.; Liu, X.; Perkins, A.
Show abstract
SummaryThe proliferation of pathogen bioinformatics pipelines has outpaced the communitys ability to compare them on common ground. Self-reported performance numbers, ad-hoc evaluation datasets, and inconsistent metrics make pipeline selection difficult for clinical and public-health researchers. We present PathoBench, an open web platform that addresses this gap through three coordinated mechanisms: (i) a curated registry of 26 standard benchmark datasets across 10 human pathogens, each with persistent identifiers and direct download links; (ii) pathogen-specific evaluation metrics that submissions must report, allowing direct head-to-head comparison only on the same dataset; and (iii) a credibility framework combining mandatory dataset attestation, ORCID-linked attribution, public peer comments, and administrator verification. As a case study, four published Mycobacterium tuberculosis drug-resistance pipelines were evaluated against the WHO TB mutation catalogue, demonstrating the frameworks discriminating power. PathoBench is open for community contributions across all ten supported pathogens. Availability and implementationPathoBench is freely available at https://pathobench.vercel.app. Source code is released under the MIT license at https://github.com/BPHL-Molecular/pathobench. The platform requires no installation for end users; programmatic access is available via a Supabase REST API. Contactyibo.dong@flhealth.gov Supplementary informationSupplementary data are available at Bioinformatics online.
Chen, Z.; Wang, R.; Luo, Q.
Show abstract
Protein language models (pLMs) offer great potential for protein sequence analysis, yet the scarcity of labeled data often limits their effectiveness in fine-tuning. Data augmentation is a promising remedy, but systematic evaluation of augmentation strategies for protein sequences remains limited, and the conditions under which augmentation confers downstream benefits are not well understood. In this paper, we systematically investigate pLM-guided substitution-based augmentation across seven protein prediction tasks. We propose ProtAug, a framework that leverages encoder-based (ESM-2) and autoregressive (ProtGPT2) pLMs to generate augmented sequences with user-controlled variation levels. Our investigation focuses on four questions: (Q1) whether pLM-synthesized sequences preserve more original signals than simpler methods, (Q2) to what extent augmentation improves prediction performance, (Q3) how variation levels affect downstream accuracy across tasks and models, and (Q4) whether biological plausibility is a necessary condition for achieving improvement. Our experimental results show that: (1) ProtAug Esm generally preserves motifs and structural similarity better than simple substitution, often comparable to homology retrieval; (2) augmentation yields consistent but task-dependent improvements, with ProtAug Esm achieving the best or second-best performance in 5 out of 7 tasks at 10% variation; (3) low-to-moderate variation levels (2-30%) perform best overall, although high-variation augmentation can benefit certain structure-related tasks; (4) the necessity of biological plausibility is task- and variation-dependent--while semantic preservation correlates with performance at low-to-moderate variation levels, improved generalization at high variation levels suggests that regularization effects, rather than label preservation, can also drive performance gains.
Chen, W.; Yang, C.; Qiu, L.; Hu, J.; Zhou, Y.
Show abstract
Summary: Long-read sequencing (LRS) has become essential for genome assembly, structural variations (SVs) detection, haplotype phasing and transcript isoform characterization. However, these applications often require manual inspection of read alignment for validation. Existing visualization tools are either interactive genome browsers that are difficult to scale to large datasets or batch-oriented tools that are not optimized for the unique alignment patterns of long-read data. We developed Bamsnap-LRS, an automated command-line tool for high-throughput LRS alignment visualization. It supports long-read-specific features, phased SNP inspection, and publication-ready batch figure generation within a unified framework for genomic, transcriptomic, and haplotype-aware analyses. Availability and Implementation: All codes and examples are freely available at https://github.com/comery/Bamsnap-LRS.
Kambara, K.; Ardie, S. W.; Tsugama, D.
Show abstract
MotivationDifferential gene expression analysis (DEA) via RNA sequencing (RNA-seq) is essential but remains challenging for wet-lab biologists due to command-line complexities. Centralized web platforms democratize this process but suffer from server congestion, long queuing delays, data privacy risks with proprietary datasets, and limited long-term sustainability due to hosting fees. ResultsWe present DEAR-OWL (Differential Expression Analysis Resource on the Web (Lite)), a fully serverless, privacy-preserving web application that performs the DEA locally inside the users web browser. To combine instant exploratory speed with rigorous verification, the application runs two distinct analysis options. The first option is a fast screening tool written in native browser language (JavaScript) that delivers immediate, genome-wide fold-change calculations and statistical screening based on an edgeR-equivalent logic. The second option is a heavy-duty statistical tool that brings the standard R package (DESeq2) directly into the browser using WebR and WebAssembly technology, ensuring publication-grade validation without needing server power. Interactive visual plots (volcano plots, minus-average plots, and heatmaps) are seamlessly generated from the results of either analysis choice. Benchmarking proved its hardware compatibility: the browser-based DESeq2 engine completed the analysis in [~]30 seconds on a 64 GB RAM workstation and in [~]3 minutes on an 8 GB RAM laptop without crashing. DEAR-OWL can utilize the Grass Expression Atlas (GExA) data as built-in and supports secure local file uploads, ensuring total data privacy with neither queuing delays nor cloud infrastructure costs. Availability and implementationDEAR-OWL is freely accessible at https://webpark2116.sakura.ne.jp/deseq2/. The source code is available at https://github.com/kota200/DEAR-OWL.
Munoz-Esquivel, G.; Fuxman Bass, J. I.; Soto-Ugaldi, L. F.
Show abstract
SummaryMapping protein regions to genomic coordinates underpins the study of exon architecture and the interpretation of clinical variants in their exon context. Existing tools resolve individual queries accurately but scale poorly to proteome-wide analyses. We present fastCDS, a C++ toolkit with command line and Python interfaces for rapid protein-to-genome coordinate mapping from GTF annotations. It matches the accuracy of existing methods while running at least two to three orders of magnitude faster. Mapping all human Pfam domains in seconds, we used the resulting atlas to examine how exonic architecture varies with domain function. Availability and ImplementationfastCDS is freely available under the MIT license at {{https://github.com/SotoLF/fastCDS}} and can be installed with pip install fastCDS or mamba install -c bioconda fastCDS. Pre-built GTF genome indices are archived at Zenodo, DOI: https://zenodo.org/records/21436146. Contactlsoto@rockefeller.edu Supplementary InformationSupplementary data are available at Bioinformatics online.