Back

Bioinformatics

Oxford University Press (OUP)

All preprints, ranked by how well they match Bioinformatics's content profile, based on 1204 papers previously published here. The average preprint has a 0.85% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
GAUSS: A comprehensive R package for accurate estimation of linkage disequilibrium for variants, Gaussian imputation and TWAS analysis of cosmopolitan cohorts

Lee, D.; Bacanu, S.-A.

2023-09-22 genetic and genomic medicine 10.1101/2023.09.19.23295783 medRxiv
Top 0.1%
86.2%
Show abstract

MotivationAs the availability of larger and more ethnically diverse reference panels grows, there is an increase in demand for ancestry-informed imputation of genome-wide association studies (GWAS), and other downstream analyses, e.g., fine-mapping. Performing such analyses at the genotype level is computationally challenging and necessitates access to individual-level genotype and phenotype data. Summary-statistics-based tools, not requiring individual-level data, provide an efficient alternative that streamlines computational requirements and promotes open science by simplifying the re-analysis and downstream analysis of existing GWAS summary data. However, existing tools perform only disparate parts of needed analysis, have only command-line interfaces and are difficult to extend/link by applied researchers. ResultsTo address these challenges, we present GAUSS -- a comprehensive and user-friendly R package designed to facilitate the re-analysis/downstream analysis of GWAS summary statistics. GAUSS offers an integrated toolkit for a range of functionalities, including i) estimating ancestry proportion of study cohorts, ii) calculating ancestry-informed linkage disequilibrium, iii) imputing summary statistics of unobserved variants, iv) conducting transcriptome-wide association studies, and v) correcting for "Winners Curse" biases. Notably, GAUSS utilizes an expansive, multi-ethnic reference panel consisting of 32,953 genomes from 29 ethnic groups. This panel enhances the range and accuracy of imputable variants, including the ability to impute summary statistics of rarer variants. As a result, GAUSS elevates the quality and applicability of existing GWAS analyses without requiring access to subject-level genotypic and phenotypic information. Availability and implementationThe GAUSS R package, complete with its source code, is readily accessible to the public via our GitHub repository at https://github.com/statsleelab/gauss. To further assist users, we provided illustrative use-case scenarios that are conveniently found at https://statsleelab.github.io/gauss/. Contactleed13@miamioh.edu Supplementary informationSupplementary data are available at Bioinformatics online.

2
Integrating Spatially Adjusted Protein Summaries for Survival Prediction in Spatial Proteomics

Ahn, S.; Oh, E. J.; Prada, D.; Shojaie, A.

2026-06-11 bioinformatics 10.64898/2026.06.08.730964 medRxiv
Top 0.1%
81.0%
Show abstract

Recent advances in spatial proteomics, particularly imaging mass cytometry, enable the measurement of protein expression at the single-cell level while preserving a spatial context. Conventional survival analyses, however, typically rely on patient-level averages of protein intensities and therefore overlook spatial heterogeneity and tissue architecture. To address this limitation, we introduce a framework that incorporates spatial information into survival modeling by generating spatially adjusted protein summaries (SAPS). In this approach, cell-level protein intensities within each patient are modeled using spatial spline regression to capture spatial trends. From these models, we extract two complementary features: a spatially adjusted mean expression and a residual variance that reflects cell-to-cell variability unexplained by spatial effects. These summaries are then incorporated into Cox proportional hazards models in combination with clinical covariates. In simulation studies, our proposed framework achieved improved predictive performance compared to other alternative methods. The application of the method to breast cancer imaging mass cytometry data indicate that spatially adjusted summaries may enhance survival prediction and reveal biologically interpretable spatial protein patterns, suggesting high translational potential. This methodology offers an efficient means of translating complex spatial proteomics data into patient-level features, providing both improved survival prediction and new insights into the role of spatial heterogeneity in cancer outcomes.

3
Yomix: An Interactive Tool for the Exploration of Low-Dimensional Embeddings in Omics Data

Perrin-Gilbert, N.; Amjad, N.; Fumeron, P.; Tulli, S.; Waterfall, J. J.

2025-12-10 bioinformatics 10.64898/2025.12.05.692343 medRxiv
Top 0.1%
77.3%
Show abstract

SummaryIn the analysis of diverse omics data, a common and important preliminary step involves computing low-dimensional embeddings using techniques such as PCA, UMAP, t-SNE, or variational autoencoders. These embeddings provide a global overview of sample distributions and their relationships, often serving as the basis for formulating biological hypotheses. To facilitate rapid and intuitive exploration of such low-dimensional embeddings, we developed Yomix, an interactive omics-agnostic visualization and data exploration tool. Yomix enables users to flexibly define subsets of interest using a lasso selection tool, instantly compute their feature signatures, and compare their distributions. Yomix is a fast and efficient tool for interactive exploration of diverse omics datasets. Availability and ImplementationYomix and its documentation are publicly available at https://github.com/perrin-isir/yomix.

4
BinDash 2.0: New MinHash Scheme Allows Ultra-fast and Accurate Genome Search and Comparisons

Zhao, J.; Zhao, X.; Pierre-Both, J.; Konstantinidis, K. T.

2024-03-14 bioinformatics 10.1101/2024.03.13.584875 medRxiv
Top 0.1%
77.0%
Show abstract

MotivationComparing large number of genomes in term of their genomic distance is becoming more and more challenging because there is an increasing number of microbial genomes deposited in public databases. Nowadays, we may need to estimate pairwise distances between millions or even billions of genomes. Few softwares can perform such comparisons efficiently. ResultsHere we update the multi-threaded software BinDash by implementing several new MinHash algorithms and computational optimization (e.g. Simple Instruction Multiple Data, SIMD) for ultra-fast and accurate genome search and comparisons at trillion scale. That is, we implemented b-bit one-permutation rolling MinHash with optimal/faster densification with SIMD. Now with BinDash 2, we can perform 0.1 trillion (or [~]10^11) pairs of genome comparisons in about 1.8 hours on a descent computer cluster or several hours on personal laptops, a [~]50% or more improvement over original version. The ANI (average nucleotide identity) estimated by BinDash is well correlated with other accurate but much slower ANI estimators such as FastANI or alignment-based ANI. In line with the findings from comparing 90K genomes ([~]10^9 comparisons) via FastANI, the 85% [~] 95% ANI gap is consistent in our study of [~]10^11 prokaryotic genome comparisons via BinDash2, which indicates fundamental ecological and evolutionary forces keeping species-like unit (e.g., > 95% ANI) together. Availability and implementationBinDash is released under the Apache 2.0 license at: https://github.com/zhaoxiaofei/bindash Contactkostas.konstantinidis@gatech.edu Supplementary informationSupplementary data are available at Bioinformatics online.

5
Predicting protein functions using positive-unlabeled ranking with ontology-based priors

Zhapa-Camacho, F.; Tang, Z.; Kulmanov, M.; Hoehndorf, R.

2024-01-31 bioinformatics 10.1101/2024.01.28.577662 medRxiv
Top 0.1%
76.9%
Show abstract

Automated protein function prediction is a crucial and widely studied problem in bioinformatics. Computationally, protein function is a multilabel classification problem where only positive samples are defined and there is a large number of unlabeled annotations. Most existing methods rely on the assumption that the unlabeled set of protein function annotations are negatives, inducing the false negative issue, where potential positive samples are trained as negatives. We introduce a novel approach named PU-GO, wherein we address function prediction as a positive-unlabeled ranking problem. We apply empirical risk minimization, i.e., we minimize the classification risk of a classifier where class priors are obtained from the Gene Ontology hierarchical structure. We show that our approach is more robust than other state-of-the-art methods on similarity-based and time-based benchmark datasets. Data and code are available at https://github.com/bio-ontology-research-group/PU-GO.

6
COCOA.jl: A Julia package for high-performance analysis of concordance and kinetic modules in biochemical networks

Schaffranke, A.; Kueken, A.; Nikoloski, Z.

2026-05-08 systems biology 10.64898/2026.05.05.722856 medRxiv
Top 0.1%
76.9%
Show abstract

SummaryRecent advances in analysis of biochemical networks have contributed the identification of their modular structure based on the concept of multi reaction dependencies and kinetic coupling of reaction rates (Kuken et al., 2022; Langary et al., 2025). Existing implementations of the algorithms to study modular structure do not scale well with the size of the networks, prohibiting their application with genome-scale networks. Here, we introduce COCOA.jl, a multithreaded Julia package for identification of concordant and kinetic modules, with applications in the study of concentration robustness. Availability and implementationCOCOA.jl is implemented in Julia 1.12.2 and is freely available under the MIT license at https://github.com/antoniofranky/COCOA.jl. It runs on Linux, macOS, and Windows; installation is supported via the Julia package manager. COCOA.jl can be called from Python via JuliaCall. Contactantonschaf@posteo.de; ankueken@uni-potsdam.de

7
AtlasLens: Metadata-centric exploration and analysis of single-cell atlases

Ashrafiyan, S.; Kosaretskii, I.; Schulz, M. H.

2026-07-10 bioinformatics 10.64898/2026.07.10.735823 medRxiv
Top 0.1%
76.5%
Show abstract

MotivationThe rapid expansion of single-cell RNA sequencing (scRNA-seq) atlases has generated datasets comprising millions of cells annotated with increasingly rich metadata, including tissue, cell type, disease status, sex, age, treatment, and temporal information. Biological questions frequently require simultaneous interrogation of multiple metadata dimensions, such as identifying specific cell populations within defined tissues, disease states, demographic groups, and time points. While existing interactive platforms facilitate visualization and analysis of scRNA-seq data, deep metadata-driven exploration and downstream analysis of atlas-scale datasets remain insufficiently supported. ResultsWe developed AtlasLens, an open-source R/Shiny application for interactive exploration of scRNA-seq datasets and integrated cellular atlases. AtlasLens enables iterative filtering across arbitrary metadata combinations, allowing users to define biologically meaningful cellular subsets and immediately perform downstream analyses. The platform integrates interactive visualization, differential expression analysis, Gene Ontology enrichment with redundancy reduction, temporal expression analysis, and context-dependent gene function profiling through GeneCOCOA. AtlasLens additionally records analysis history and automatically generates corresponding R code to enhance reproducibility. The application is distributed through Docker for simple local deployment, preserving data privacy and eliminating dependency-management challenges. We demonstrate AtlasLens using the Tabula Muris and a time-resolved whole-lung single-cell atlas of bleomycin-induced lung injury and fibrosis, highlighting its ability to support complex metadata-driven biological investigations. AvailabilitySource code is available at https://github.com/SchulzLab/AtlasLens. Contactmarcel.schulz@em.uni-frankfurt.de

8
Corpus-wide causality: Algorithm design & application for aggregating gene-disease causal evidence

Bansal, N.; Parsodkar, A. P.; Pathak, A.; Narayanan, M.

2026-05-12 bioinformatics 10.64898/2026.05.08.723796 medRxiv
Top 0.1%
76.0%
Show abstract

Identifying causal relationships and distinguishing them from associations is a central scientific endeavor with many applications; knowing causal links between genes and diseases, for instance, can focus drug discovery on curing diseases beyond just symptom management. Despite several studies on automatically extracting relations between entities from large biomedical literature corpora like PubMed, only a few studies extract causal relations from abstracts and even fewer summarize corpus-level evidence for causal links. Recently, Large Language Models (LLMs) have been increasingly deployed to summarize biomedical information and extract relations; however, there is a distinct lack of explicit benchmarking comparing these generalized LLM-based methods against specialized, domain-aware frameworks for corpus-wide causal inference. In this work, we develop a method to infer Corpus-Wide Causal Score (CWCS) of a gene-disease (G-D) pair by integrating two pieces of evidence: (i) network-based causal signals in a prior gene regulatory network, quantified as a CWCS-Net score using an existing multilayer network centrality algorithm; and (ii) corpus-wide literature evidence, quantified as a CWCS-TD (TD for Truth Discovery) score using a newly-developed TD algorithm. Our CWCS-TD (scoring) algorithm jointly and iteratively estimates causal scores for multiple G-D pairs while modeling the reliability of PubMed abstracts co-mentioning them; and represents an advance in the field of TD algorithms due to its incorporation of bibliometric features of publications to address the challenge of sparsity of abstracts that assert a G-D causal relation. Using OMIM as an external expert-curated reference to evaluate classifications of G-D pairs as causal or not, our CWCS method achieved a causal class F1 score of 0.600 across ten diseases, outperforming both LLMs, GPT-4o and MMed-Llama 3 (this performance trend also persists when using area under the precision-recall curve as the evaluation metric). Both LLMs exhibit high recall accompanied by comparatively low precision, resulting in lower causal class F1 scores (0.505 for GPT-4o and 0.522 for MMed-Llama 3) due to large number of false positive predictions. Taken together, these evaluations and other ablation studies show the promise of our carefully designed algorithm in collating and integrating evidence of biomedical causal relations from both network- and literature-based sources, thereby supporting its broader applicability.

9
Self-normalizing learning on biomedical ontologies using a deep Siamese neural network

Smaili, F. Z.; Gao, X.; Hoehndorf, R.

2020-04-25 bioinformatics 10.1101/2020.04.23.057117 medRxiv
Top 0.1%
75.6%
Show abstract

MotivationOntologies are widely used in biomedicine for the annotation and standardization of data. One of the main roles of ontologies is to provide structured background knowledge within a domain as well as a set of labels, synonyms, and definitions for the classes within a domain. The two types of information provided by ontologies have been extensively exploited in natural language processing and machine learning applications. However, they are commonly used separately, and thus it is unknown if joining the two sources of information can further benefit data analysis tasks. ResultsWe developed a novel method that applies named entity recognition and normalization methods on texts to connect the structured information in biomedical ontologies with the information contained in natural language. We apply this normalization both to literature and to the natural language information contained within ontologies themselves. The normalized ontologies and text are then used to generate embeddings, and relations between entities are predicted using a deep Siamese neural network model that takes these embeddings as input. We demonstrate that our novel embedding and prediction method using self-normalized biomedical ontologies significantly outperforms the state-of-the-art methods in embedding ontologies on two benchmark tasks: prediction of interactions between proteins and prediction of gene-disease associations. Our method also allows us to apply ontology-based annotations and axioms to the prediction of toxicological effects of chemicals where our method shows superior performance. Our method is generic and can be applied in scenarios where ontologies consisting of both structured information and natural language labels or synonyms are used. Availabilityhttps://github.com/bio-ontology-research-group/Ontology-based-normalization Contactrobert.hoehndorf@kaust.edu.sa and xin.gao@kaust.edu.sa

10
AtlasXbrowser enables spatial multi-omics data analysis through the precise determination of the region of interest

Barnett, J.; Silverman, J.; Wetzel, M.; Rao, P.; Sotudeh, N.; Wang, L.

2022-06-02 bioinformatics 10.1101/2022.05.11.491526 medRxiv
Top 0.1%
75.5%
Show abstract

Recent developments in novel spatial sequencing technologies allow for the incorporation of spatial information into high-throughput sequencing assays. One such method, Deterministic Barcoding in Tissue for spatial omics sequencing (DBiT-seq, abbreviated herein as DBiT), utilizes perpendicular microfluidic channels to deliver DNA barcodes across the tissue in a spatially-encoded manner, allowing for sequenced reads to be mapped back onto the 2-D coordinates of the tissue to provide spatial coordinates to cells. DBiT has been the first spatial sequencing technology developed for epigenomic assays beyond transcriptome and proteome. However, despite existing of many open-source software packages for downstream bioinformatics analysis, there is no software available for processing DBiT image data with evenly spaced channels. To facilitate the integration of DBiT spatial and sequenced data, here we proposed a new method to precisely capture the spatial information and further developed AtlasXbrowser based on the new method to extract spatial data from the image data. AtlasXbrowser is a python-based tool with GUI that requires no technical expertise to operate and enables researchers to incorporate brightfield and epifluorescence images of processed tissue samples into downstream bioinformatics analysis tools. Availability and implementationFreely available at https://github.com/atlasxomics/AtlasXbrowser.

11
DeepGOZero: Improving protein function prediction from sequence and zero-shot learning based on ontology axioms

Kulmanov, M.; Hoehndorf, R.

2022-01-14 bioinformatics 10.1101/2022.01.14.476325 medRxiv
Top 0.1%
75.4%
Show abstract

MotivationProtein functions are often described using the Gene Ontology (GO) which is an ontology consisting of over 50,000 classes and a large set of formal axioms. Predicting the functions of proteins is one of the key challenges in computational biology and a variety of machine learning methods have been developed for this purpose. However, these methods usually require significant amount of training data and cannot make predictions for GO classes which have only few or no experimental annotations. ResultsWe developed DeepGOZero, a machine learning model which improves predictions for functions with no or only a small number of annotations. To achieve this goal, we rely on a model-theoretic approach for learning ontology embeddings and combine it with neural networks for protein function prediction. DeepGOZero can exploit formal axioms in the GO to make zero-shot predictions, i.e., predict protein functions even if not a single protein in the training phase was associated with that function. Furthermore, the zero-shot prediction method employed by DeepGOZero is generic and can be applied whenever associations with ontology classes need to be predicted. Availabilityhttp://github.com/bio-ontology-research-group/deepgozero Contactrobert.hoehndorf@kaust.edu.sa

12
dynUGENE: an R package for uncertainty-aware gene regulatory network inference, simulation, and visualization

Lu, T.; Silva, A.

2021-01-08 bioinformatics 10.1101/2021.01.07.425782 medRxiv
Top 0.1%
75.0%
Show abstract

Methods for gene regulatory network inference focus on network architecture identification but neglect model selection and simulation. We implement an extension to the dynGENIE3 algorithm that accounts for model uncertainty as an R package, providing users with an easy to use interface for model selection and gene expression profile simulation. Source code is available at https://github.com/tianyu-lu/dynUGENE with a detailed user guide. A webserver with interactive controls is available at https://tianyulu.shinyapps.io/dynUGENE/.

13
AntiDIF: Accurate and Diverse Antibody Specific Inverse Folding with Discrete Diffusion

Branson, N.; Deane, C.

2025-07-17 immunology 10.1101/2025.07.12.664553 medRxiv
Top 0.1%
74.6%
Show abstract

Inverse folding is an important step in current computational antibody design. Recently deep learning methods have made impressive progress in improving the sequence recovery of antibodies given their 3D backbone structure. However, inverse folding is often a one-to-many problem, i.e. there are multiple sequences that fold into the same structure. Previous methods have not taken into account the diversity between the predicted sequences for a given structure. Here we create AntiDIF an Antibody-specific discrete Diffusion model for Inverse Folding. Compared with stateof-the-art methods we show that AntiDIF improves diversity between predictions while keeping high sequence recovery rates. Furthermore, forward folding of the generated sequences shows good agreement with the target 3D structure.

14
An inductive, supervised approach for predicting gene--disease associations using phenotype ontologies

Bakheet, S.; Zhapa-Camacho, F.; Hoehndorf, R.

2025-05-13 bioinformatics 10.1101/2025.05.07.652682 medRxiv
Top 0.1%
74.3%
Show abstract

MotivationPredicting gene-disease associations (GDAs) is the problem to determine which gene is associated with a disease. The problem can be framed as a ranking problem where genes are ranked based on a set of phenotypes using a measure of phenotype similarity. When phenotypes are described using phenotype ontologies, ontology-based semantic similarity measures are used. Traditional semantic similarity measures use only the ontology taxonomy. Recent methods based on ontology embeddings compare phenotypes in latent space; these methods can use all ontology axioms as well as a supervised signal, but are inherently transductive, i.e., queries must already be known at the time of learning embeddings, and therefore these methods do not generalize to novel diseases (sets of phenotypes) at inference time. ResultsWe developed an inductive method for ranking genes based on a set of phenotypes. Our method first uses a graph projection to map axioms from phenotype ontologies to a graph structure, and then uses ontology embeddings to create latent representations of phenotypes. We use an explicit aggregation strategy to combine phenotype embeddings into representations of genes or diseases, allowing us to generalize to novel sets of phenotypes. We also develop a method to make the phenotype embeddings and the similarity measure task-specific by including a supervised signal from known gene-disease associations. We apply our method to mouse models of human disease and demonstrate that we can significantly improve over inductive baseline measures, and reach a performance similar to transductive methods for predicting gene-disease associations while being more general. Availability and Implementationhttps://github.com/bio-ontology-research-group/ISGDA Contactrobert.hoehndorf@kaust.edu.sa

15
A joint embedding of protein sequence and structure enables robust variant effect predictions

Blaabjerg, L. M.; Jonsson, N.; Boomsma, W.; Stein, A.; Lindorff-Larsen, K.

2023-12-16 bioinformatics Community evaluation 10.1101/2023.12.14.571755 medRxiv
Top 0.1%
74.1%
Show abstract

The ability to predict how amino acid changes may affect protein function has a wide range of applications including in disease variant classification and protein engineering. Many existing methods focus on learning from patterns found in either protein sequences or protein structures. Here, we present a method for integrating information from protein sequences and structures in a single model that we term SSEmb (Sequence Structure Embedding). SSEmb combines a graph representation for the protein structure with a transformer model for processing multiple sequence alignments, and we show that by integrating both types of information we obtain a variant effect prediction model that is more robust to cases where sequence information is scarce. Furthermore, we find that SSEmb learns embeddings of the sequence and structural properties that are useful for other downstream tasks. We exemplify this by training a downstream model to predict protein-protein binding sites at high accuracy using only the SSEmb embeddings as input. We envisage that SSEmb may be useful both for zero-shot predictions of variant effects and as a representation for predicting protein properties that depend on protein sequence and structure.

16
Omix: A Multi-Omics Integration Pipeline

Schneegans, E.; Fancy, N.; Thomas, M.; Willumsen, N.; Matthews, P. M.; Jackson, J.

2023-09-01 bioinformatics 10.1101/2023.08.30.555486 medRxiv
Top 0.1%
73.6%
Show abstract

SummaryThe Omix pipeline offers an integration and analysis framework for multiomics intended to preprocess, analyse, and visualise multimodal data flexibly to address various research questions. From biomarker discovery and patient stratification to the investigation of complex biological processes, Omix empowers researchers to derive valuable insights from omics data. Using Alzheimers Disease (AD) bulk proteomics and transcriptomics datasets generated from two distinct regions derived from post-mortem brains, we demonstrate the utility of Omix in generating an integrated pseudo-temporal multi-omics profile of AD. Availability and ImplementationOmix is implemented as a software package in R. The code for the Omix package is available at https://github.com/eleonoreschneeg/Omix. Reference documentation and online tutorials are available at https://eleonore-schneeg.github.io/Omix. All code is open-source and available under the GNU General Public License v3.0 (GPL-3). Contacteleonore.schneegans17@imperial.ac.uk, johanna.jackson@imperial.ac.uk

17
ChemoCalib: multiblock PLS calibration of genome-scale metabolic models improves flux prediction over expression-only integration

zhang, X.

2026-07-31 bioinformatics 10.64898/2026.07.28.741216 medRxiv
Top 0.1%
73.5%
Show abstract

MotivationConstraint-based metabolic modeling faces a calibration gap: genome-scale metabolic models (GEMs) integrated with transcriptomics alone rely on expression-to-flux heuristics (E-Flux, GIMME, iMAT, MOMENT) that ignore cross-omics co-variance structure and lack statistical mechanisms for propagating omics uncertainty into reaction bounds, yielding flux predictions with limited agreement to 13C metabolic flux analysis (MFA) measurements. ResultsWe present Chemo-Calib, a multiblock PLS (MB-PLS) framework that calibrates GEM reaction bounds from the shared latent structure of metabolomics, transcriptomics, and proteomics data. On 11 E. coli 13C-MFA reference conditions spanning the Keio fluxome and Holm 2010 datasets, ChemoCalib constrained FBA on iJO1366 achieves a Spearman{rho} = 0.461 overall (up to 0.523 in PPP) and Pearson r of 0.49-0.58 across central carbon pathways, with statistically significant improvement over expression-only baselines including E-Flux2 and SPOT (p < 0.05, Holm-corrected). The latent-to-constraint mapping employs GPR-aware VIP aggregation (Algorithm 1) to project multi-omics latent scores onto genome-scale reaction bounds without heuristic thresholding. An optional in-silico active learning loop (relegated to Supplementary Material) further tightens calibration through virtual experiment selection. AvailabilityChemoCalib is open-source (MIT) at https://github.com/chemocalib/chemocalib with Docker support, a 5-minute tutorial, and pre-computed iJO1366 benchmarks. Preprint available at bioRxiv; code archived at Zenodo DOI: 10.5281/zenodo.21645890.

18
BioRSP: a method for characterizing enrichment patterns in single-cell embeddings

Yao, Z.; Chen, J. Y.

2025-12-29 bioinformatics 10.1101/2024.06.25.599250 medRxiv
Top 0.1%
73.0%
Show abstract

Low-dimensional embeddings such as UMAP and t-SNE are routinely used to visually interpret high-dimensional omics data, yet claims based on embedding geometry are often qualitative, embedding-sensitive, and weakly calibrated. We present BioRSP (Biological Radar Scanning Plots), a geometry-first framework that quantifies how a user-defined foreground subset is distributed across the embedding footprint of a fixed analysis set. BioRSP converts 2D coordinates to polar form around a robust vantage point, scans the set by angle, and computes a radar profile that summarizes signed radial enrichment of foreground relative to background within sliding angular windows using a distance-based radial discrepancy. The profile is reduced to interpretable summaries including anisotropy magnitude, peak directionality, and coverage, and is accompanied by explicit adequacy rules and subsampling-based stability diagnostics so the method can abstain when the geometry is underpowered. We demonstrate BioRSP in a community-standard human kidney single-nucleus reference by analyzing thick ascending limb (TAL) nuclei using published UMAP coordinates and a standardized within-set top-decile foreground rule. Within TAL, BioRSP distinguishes sharply localized rim-enrichment patterns, broadly supported but structured within-type heterogeneity, and near-isotropic profiles, with anisotropy spanning more than an order of magnitude in the pooled TAL analysis. Donor-aware reruns show that per-donor adequacy is frequently limiting, but that directional profiles are stable when donor-level support is sufficient. BioRSP is provided as open-source software producing standardized plots, summary tables, and run manifests to support reproducible embedding-aligned enrichment analysis.

19
ABDS: tool suite for analyzing biologically diverse samples

Du, D.; Bhardwaj, S.; Parker, S. J.; Cheng, Z.; Zhang, Z.; Lu, Y.; Van Eyk, J. E.; Yu, G.; Clarke, R.; Herrington, D. M.; Wang, Y.

2023-07-05 bioinformatics 10.1101/2023.07.05.547797 medRxiv
Top 0.1%
72.6%
Show abstract

MotivationAnalytics tools are essential to identify informative molecular features about different phenotypic groups. Among the most fundamental tasks are missing value imputation, signature gene detection, and expression pattern visualization. However, most commonly used analytics tools may be problematic for characterizing biologically diverse samples when either signature genes possess uneven missing rates across different groups yet involving complex missing mechanisms, or multiple biological groups are simultaneously compared and visualized. ResultsWe develop ABDS tool suite tailored specifically to analyzing biologically diverse samples. Mechanism-integrated group-wise imputation is developed to recruit signature genes involving informative missingness, cosine-based one-sample test is extended to detect enumerated signature genes, and unified heatmap is designed to comparably display complex expression patterns. We discuss the methodological principles and demonstrate the conceptual advantages of the three software tools. We also showcase the biomedical applications of these individual tools. Implemented in open-source R scripts, ABDS tool suite complements rather than replaces the existing tools and will allow biologists to more accurately detect interpretable molecular signals among diverse phenotypic samples. Availability and implementationThe R Scripts of ABDS tool suite is freely available at https://github.com/niccolodpdu/ABDS. Contactyuewang@vt.edu Supplementary informationSupplementary materials are available at Bioinformatics Advances online.

20
TransporterPAL: An integrative database Transporter Prediction ALgorithm

Dyekjaer, J. D.; Andersen, A. K. B.; Madsen, J. A. V.; Morth, J. P.; Jensen, L. J.; Borodina, I.

2022-09-06 bioinformatics 10.1101/2022.09.05.506577 medRxiv
Top 0.1%
72.6%
Show abstract

MotivationNatural products are used as drugs, cosmetic ingredients, pigments, flavors, and agricultural products. The compounds are retrievable as extracts from natural sources, but the yields are often low, and the final product may contain various impurities. These challenges can be solved by expressing the biosynthetic pathway in microbial cell factories and ensuring product secretion from the cell by using an appropriate transporter. However, insufficient knowledge of transporters for specific compounds often obstructs efficient secretion of the natural product. Therefore, our goal was to develop an algorithm that predicts transporters for a given compound using available public data. ResultsThe web application TransporterPAL predicts suitable transporters for compounds by interconnecting data for biosynthetic genes and their interactions with transporters. The web application queries the STITCH, STRING, and UniProtKB databases via their respective APIs and returns a set of potential transporters based on a compound and, optionally, the organism as input. For a test set of 61 transporter systems, each containing one or more transporters, a total of 90 unique transporters with a known substrate, we could retrieve 45% of the transporters. To our knowledge, this is the first bioinformatics tool for predicting transporter candidates for a given molecule. Availabilityhttps://transporterpal.com Contactirbo@biosustain.dtu.dk Supplementary informationSupplementary data are available at Bioinformatics online.