Back

SoftwareX

Elsevier BV

Preprints posted in the last 90 days, ranked by how well they match SoftwareX's content profile, based on 15 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
A Python Dash App and cPanel workflow to automate metabolomics data analyses and visualisation

O'Loughlin, J.; Moses, T.

2026-05-05 biochemistry 10.64898/2026.05.01.722139 medRxiv
Top 0.1%
18.6%
Show abstract

Metabolomics offers a sophisticated analytical framework for characterising the molecular phenotype of biological organisms and complex living systems at a high resolution. As the functional endpoint of the omics cascade, the metabolome serves as a close reflection of cellular activity. It integrates genetic, transcriptomic and proteomic variations with external environmental influences. However, the inherent complexity of metabolomic datasets, characterised by high-dimensional chemical diversity, wide dynamic ranges, and significant matrix effects, necessitates a rigorous suite of chemometric and bioinformatic workflows. For researchers uninitiated in computational biology, the multi-stage requirement for raw data pre-processing, signal deconvolution, and multivariate statistical modelling (such as PCA or PLS-DA) presents a substantial barrier to entry. Navigating these convoluted data architectures remains a primary challenge in deriving biological meaning from the global metabolic profile. Here, we present a workflow to use Python Dash Apps to create a user-friendly interface for simplifying data processing and statistical calculations. Users can select their desired samples to initiate calculations for various statistical tests, generating interactive and publication-quality figures to explore their results. These apps were deployed on an Apache server via cPanel, allowing individuals to share their findings with collaborators and for research facilities to share metabolomics results with their users.

2
AnimalTA: A simple yet flexible tool for video tracking and manual corrections.

Chiara, V.; Buatois, A.; Kim, S.-Y.

2026-06-30 animal behavior and cognition 10.64898/2026.06.27.733780 medRxiv
Top 0.1%
12.7%
Show abstract

1. Video-tracking programs have now become an essential tool for researchers measuring animal behavior across biological fields. The panel of available programs is growing rapidly, providing researchers with numerous specific tools that will match their precise needs. However, their proliferation may complicate post-tracking data processing, and some programs do not even provide tools for correcting tracking errors or analysing tracking data. In the case of commercial software, the loss of access to a program due to budget limitations or researchers' mobility from one institution to another could prevent them from accessing and visualizing their tracking data. 2. There is therefore a growing need for an accessible and flexible tool to handle post-tracking processes such as the correction and analysis of tracking data obtained across different video-tracking programs. 3. We present here the latest update of the video tracking and analysis program AnimalTA. With this new release, we propose to solve the above-mentioned problems by providing the scientific community with a program that will allow for data importation from other video-tracking programs. Like in its previous versions, AnimalTA remains a free, open-source, and highly user-friendly program, ensuring that it will always be accessible without restriction. Now, with this new importation option, users who performed their tracking with other programs can benefit from AnimalTA's complete toolset of data visualization, correction, and analysis. 4. Finally, this article gives an overview of the other main improvements associated with this new release. The program is now faster in both video importation and tracking, proposes an amplified toolset for data visualisation and correction, and features new options for data analysis.

3
CcpNmr AnalysisDynamics: a unified framework for NMR dynamics data analysis

Mureddu, L. G.; Brooksbank, E. J.; Vuister, G. W.; Muskett, F. W.

2026-06-20 biochemistry 10.64898/2026.06.19.733360 medRxiv
Top 0.1%
7.2%
Show abstract

Nuclear Magnetic Resonance (NMR) relaxation experiments provide a powerful residue-resolved access to biomolecular dynamics across a wide range of timescales. Unfortunately, the quantitative analysis of the relaxation data remains distributed across specialised and often disconnected tools. Here, we present CcpNmr AnalysisDynamics, the latest addition to the CcpNmr Analysis program suite, providing an integrated framework for relaxation analysis, exchange-aware interpretation and dynamical modelling. The platform unifies relaxation-rate extraction, diagnostic validation, model-based analysis and structural visualisation within reproducible workflows, while supporting future extension through a robust application programming interface and plugin architecture. We introduce ModelAnalysis (ModA), a new analysis engine based on the Lipari-Szabo formalism that incorporates robust optimisation, uncertainty estimation and model-selection strategies designed for heterogeneous relaxation datasets. The framework also supports exchange-focused analysis and integration with specialised external modelling tools, allowing relaxation anomalies to be followed from initial detection to more detailed interpretation. The applicability and reliability of AnalysisDynamics are demonstrated through systematic re-analysis and validation of curated relaxation datasets from the Biological Magnetic Resonance Data Bank. These analyses enable assessment of data consistency, dynamic parameters and model reliability across magnetic fields, providing a reproducible route from NMR relaxation measurements to structure-linked interpretation of biomolecular dynamics.

4
The NMR Exchange Format (NEF): Specification and Applications

Ploskon, E.; Baskaran, K.; Tejero, R.; Schwieters, C. D.; Bardiaux, B.; Guentert, P.; Fogh, R. H.; Gutmanas, A.; Brooksbank, E. J.; Yokochi, M.; Wishart, D. S.; Wedell, J. R.; Vranken, W. F.; Thompson, D.; Thompson, G.; Smith, B. O.; Rehman, S.; Ramelot, T. A.; Ragan, T. J.; Perez, A.; Perera, B. L.; Peisach, E.; Nilges, M.; Mureddu, L. G.; Mondal, A.; Lubicka, E. A.; Liwo, A.; Kurisu, G.; Kobayashi, N.; Klukowski, P.; Johnston, B. A.; Huang, Y. J.; Hoch, J. C.; Higman, V. A.; Herrmann, T.; Hayward, M. W.; Garnet, J. A.; Case, D. A.; Burley, S. K.; Adams, P. D.; Montelione, G. T.; Vuister, G.

2026-04-24 biochemistry 10.64898/2026.04.22.715536 medRxiv
Top 0.1%
5.5%
Show abstract

The NMR Exchange Format (NEF) is a community-driven standard for representing NMR experimental data in a consistent, interoperable, and machine-readable form. Built on the STAR syntax, NEF provides a structured framework for storing and exchanging chemical shifts, peak lists, various types of structural restraints, and related metadata, thus allowing for data exchange across software platforms. By enabling direct, lossless transfer of information, NEF simplifies multi-software workflows, improves reproducibility, and supports FAIR (Findable, Accessible, Interoperable, Reusable) data principles. We describe the NEF specification, its current implementation across commonly used NMR software packages, and its application in areas including biomolecular structure determination, metabolomics, and ligand screening. Testing demonstrates that NEF can be used to exchange complete datasets between programs without loss of information or functionality. We also outline recent developments and future directions, such as inclusion of NMR relaxation data and support for non-standard residue topologies. NEFs growing adoption highlights its potential as a unifying standard for NMR data, enabling more efficient, transparent and collaborative research.

5
Neurokraken: A fully flexible, open-source, python-based neuroscience behavior platform

Wallerus, A.; Castro e Almeida, S.; Passecker, J.

2026-07-06 animal behavior and cognition 10.64898/2026.06.30.735592 medRxiv
Top 0.1%
5.3%
Show abstract

A major challenge in behavioral neuroscience is the lack of a unified software framework capable of implementing diverse paradigms across species and experimental setups. Researchers currently face a trade-off: they must either spend significant time developing custom, siloed solutions that hinder reproducibility, or incur substantial costs purchasing inflexible, closed systems. Here, we present Neurokraken, an open-source, Python-native platform designed to overcome these limitations. Neurokraken allows writing experiment progression entirely in standard python, while its core architecture automatically sets up a microcontroller for the connected hardware components and enables python side access with millisecond-precision timing and automatic logging. The system prioritizes ease of use and flexibility, enabling advanced series of events and conditions, the usage of python ecosystem code and packages within experiments, and the addition of any arduino-compatible electronic devices for custom experiments. As a result, users can easily create interactive virtual and real environments to engage, monitor, and record subjects. We present Neurokraken's versatility across a wide range of paradigms, for human and non-human primate psychophysics, and complex rodent behavior in both head-fixed and freely moving paradigms. Its modular design allows for rapid hardware reconfiguration, while a fully customizable user interface enables real-time monitoring and interactive experimental control without compromising timing precision. By uniting laboratory-grade precision with an accessible and flexible open-source philosophy, Neurokraken provides a single, powerful solution to design and execute next-generation behavioral experiments. We hope Neurokraken helps accelerate research, improve reproducibility throughout the neuroscience community, and make advanced behavioral experimentation more accessible through its substantial cost-efficiency.

6
OMIO: A policy-driven Python library for reproducible microscopy image I/O

Musacchio, F.; Antony, H.; Crux, S.; Fuhrmann, F.; Gockel, N.; Hoffmann, D. M.; Mercan, D.; Nebeling, F. C.; Fuhrmann, M.

2026-06-11 bioinformatics 10.64898/2026.06.09.731118 medRxiv
Top 0.1%
3.3%
Show abstract

Modern fluorescence and multiphoton microscopy workflows operate within a heterogeneous ecosystem of file formats, partially overlapping metadata standards, and reader-specific conventions. In practice, this frequently leads to silent axis misinterpretations, loss or corruption of physical voxel size information, and laboratory-specific glue code that is fragile, poorly documented, and difficult to reproduce. OMIO, short for Open Microscopy Image I/O, addresses these issues by providing a lightweight, policy-driven image I/O layer for Python that enforces a canonical, OME-compatible data representation at the API boundary. The central contribution of OMIO is the explicit separation of low-level format access from semantic normalization. Existing reader libraries are used as interchangeable backends for extracting pixel data and available metadata, while OMIO enforces axis conventions, metadata interpretation, and fallback decisions in a centralized and auditable policy layer. This design allows heterogeneous microscopy inputs to be converted into a stable representation without propagating backend-specific assumptions into downstream analysis code. The core design principles of OMIO include canonical axis semantics (TZCYX), robust metadata normalization with explicit and auditable fallbacks, memory-aware operation via optional Zarr-based backends, and workflow-level semantics that extend beyond individual files to folder stacks and BIDS-like project structures. This architecture allows OMIO to orchestrate existing reader libraries into a coherent and reproducible I/O pipeline without replacing or duplicating their functionality. OMIO is implemented as an open-source and community-oriented system in which support for additional file formats and metadata conventions can be added incrementally through modular reader backends. By encouraging the contribution of example datasets, backend extensions, and feature requests, OMIO is designed to evolve alongside emerging acquisition systems while preserving strict semantic guarantees at the interface level. The resulting standardized OME-TIFF outputs are immediately suitable for downstream quantitative analysis and interactive inspection in scientific Python workflows, including workflows based on ImageJ and Napari.

7
HydraMPP: A lightweight library for distributed massive parallel processing in Python - threading at scale.

Figueroa, J. L.; White, R. A.

2026-06-08 bioinformatics 10.64898/2026.06.04.730204 medRxiv
Top 0.1%
2.4%
Show abstract

We now exist in the era of massive datasets from genomics, large language models, and all the known knowledge of humanity right at our fingertips. Much of this data is becoming more accessible; however, processing such data remains an ongoing issue across systems including high performance computing (HPC) infrastructures. Massively parallel computing (MPP) has solved this using a divide and conquer approach by splitting workloads across independent nodes (i.e., central processing units (CPU) allowing for higher scaling of data). The main engine for this in python is Ray; however, it has many issues including a large code space, security issues, debugging opacity, and memory management issues. Here, we present HydraMPP, a lightweight, ease of use and utilization, with high auditability, and with SLURM ergonomics.

8
MicrobeMS - A MATLAB Toolbox for Microbial Identification Based on Mass Spectrometry

Lasch, P.

2026-05-12 bioinformatics 10.64898/2026.05.08.723807 medRxiv
Top 0.1%
2.4%
Show abstract

1.Over the last two decades, matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-ToF MS) has become the standard method for identifying bacteria and has found a wide range of applications, especially in clinical microbiology. The methods high taxonomic resolution, minimal sample preparation, and complete, ready-to-use commercial systems, which include instrumentation, experimental protocols, spectral databases, and identification analysis software, were key factors in the success of MALDI-ToF MS as the standard for identifying microorganisms in routine diagnostic laboratories. However, despite the availability of these commercial solutions, there is also a growing need for efficient, cost-effective, vendor-neutral databases and analysis tools. These tools would enable the compilation of user-defined mass spectral databases and the testing of new analysis methods and algorithms, particularly in an academic context. To this end, MicrobeMS software has been developed to cover all stages of MALDI-ToF MS-based identification analysis. MicrobeMS is an easy-to-use desktop application for analyzing mass spectra from microorganisms and performing tasks related to spectrum database compilation. It includes routines for direct data import and export, biomarker peak searches, management of spectrum metadata, testing of spectrum quality, supervised and unsupervised identification analysis and intuitive result display. MicrobeMS is implemented in MATLAB and is freely available as MATLAB pcode for Windows and Linux, as well as a standalone application. Over the last fifteen years, the software has undergone continuous development and is now used routinely in various settings at the Centre for Biological Threats and Special Pathogens (ZBS) at the Robert Koch Institute (RKI) in Berlin, Germany, for example in supporting spectrum database compilation, to identify special or rare pathogenic bacteria by advanced identification analysis concepts, or to test in silico MALDI-ToF MS databases derived from microbial genomes. In this software publication the versatility and capabilities of MicrobeMS are demonstrated using a test data set from highly pathogenic bacteria (HPB) which has been obtained as part of a published European Union (EU)-funded External Quality Assurance Exercise (EQAE). MicrobeMS and HPB test data can both be downloaded from https://wiki.microbe-ms.com/. The goal of this software publication is twofold: to raise awareness of MicrobeMS within the scientific community and to encourage the testing of the software and custom-developed MALDI-ToF MS databases of the RKI, which are published at the ZENODO data repository (https://doi.org/10.5281/zenodo.7702374).

9
BaSiCPy: Scalable and Robust Shading Correction for Optical Microscopy Images

Liu, Y.; Fukai, Y. T.; Cano-Muniz, S.; Perez, V.; Todorov, M.; Ortega, G.; Morello, T.; Loeffler, D.; Paetzold, J.; Xu, X.; Lamm, L.; Ma, N.; Erturk, A.; Schroeder, T.; Boeck, L.; Schapiro, D.; Schaub, N.; Marr, C.; Peng, T.

2026-05-01 bioengineering 10.64898/2026.04.28.721386 medRxiv
Top 0.1%
2.4%
Show abstract

Quantitative fluorescence microscopy is frequently confounded by spatially varying illumination and temporal intensity drift. Although BaSiC is a widely adopted retrospective correction method, it can fail when foreground content is strongly correlated across images--a common regime in time-lapse, tiled and volumetric acquisitions--and its application often requires manual parameter tuning that limits reproducibility and scalability. We introduce BaSiCPy, a foreground-aware implementation of BaSiC that improves illumination profile estimation under correlated foreground structures, provides automatic hyperparameter selection and accelerates large-scale processing through GPU support. BaSiCPy is distributed as an open-source Python package with graphical and programmatic interfaces, facilitating integration into contemporary bioimage analysis workflows.

10
CICADA: A unified framework for NWB-based neurophysiological data analysis

Hamon, M.; Lebert, J.; Denis, J.; Filippi, C.; Renard, A.; Bech, P.; Pulin, M.; Bisi, A.; Molinuevo Gomez, D.; Priestley, J. B.; Crochet, S.; Petersen, C. C.; Cossart, R.; Picardo, M. A.; Dard, R. F.

2026-07-08 neuroscience 10.64898/2026.07.03.736318 medRxiv
Top 0.1%
2.4%
Show abstract

Neurophysiology datasets are becoming increasingly complex, combining behavioral measurements with high-dimensional neuronal activity recordings coming from optical and/or electrophysiological measurements. The Neurodata Without Borders (NWB) standard has emerged in the community as the format of record. While standardized and widely used preprocessing tools generating NWB files have been developed, extensible frameworks for scientific analysis downstream of the NWB ecosystem are still under-represented. We present CICADA, a Python framework dedicated to analysis of neurophysiological data in the standardized NWB format. The toolbox is built as three hierarchically-organized packages: cicada-nwb (NWB access layer), cicada-analysis (plugin-based analysis engine and tool library), and cicada-gui (PyQt5 desktop application at the head of the pipeline). Beyond this architectural separation, CICADA is built around a central design principle: supporting a continuum from turnkey use to full modularity. Researchers can use the complete GUI-driven cicada-gui workflow without writing code, programmatically use existing analysis plugins from cicada-analysis, contribute to new analysis plugins, reuse utilities from cicada-tools, or build entirely custom pipelines on top of the cicada-nwb access layer alone. The same analysis plugin runs identically in interactive GUI and parameter-configured headless modes, enabling reproducible multi-session, multi-animal group analyses. We illustrate the versatility of CICADA with example analyses of behavioral, calcium imaging (two-photon and widefield) and extracellular electrophysiology datasets from rodent laboratories. CICADA is open source, actively maintained, and designed so that any laboratory can contribute at any level of the stack without modifying the core framework.

11
msaGUI: Multispectral Analysis Graphical User Interface for Ratiometric Analysis and Background Correction

Hoy, G. R.; Davis, C. M.

2026-07-03 biophysics 10.64898/2026.06.30.735666 medRxiv
Top 0.1%
1.8%
Show abstract

Chemical imaging is a powerful branch of modern microscopy encumbered by a lack of flexible, high-throughput analysis tools. Bespoke analytical pipelines typically perform ratiometric analysis on two layers in a multispectral image to describe the relative composition of molecules in a sample. This strategy has been implemented across fields, spanning histopathology, cell biology, environmental science, and materials science. The commercialization of chemical imaging microscopes has facilitated the collection of large multispectral datasets, necessitating accessible ways to process them. This paper describes Multispectral Analysis Graphical User Interface (msaGUI), a desktop graphical user interface to analyze individual and batch datasets of multispectral images. Data is loaded as CSV, TSV, or TIFFs and processed through a user-defined sequence of modular image operations that can be flexibly combined, e.g. to reduce spectral crosstalk or background noise. After analysis, data is visualized as exportable images, histograms, and statistics. To yield publication-quality figures, outputted images are fully customizable. Written in Python with open-source libraries, the msaGUI program is packaged into an executable for Windows and Mac for a fully no-code application. Other operating systems are supported via the Python source code. In summary, msaGUI provides a rapid and user-friendly solution for analyzing and visualizing multispectral data.

12
PALMS: A Computational Implementation for Pavlovian Associative Learning Models Simulation

Fixman, M.; Abati, A.; Jimenez Nimo, J.; Lim, S.; Mondragon, E.

2026-05-08 animal behavior and cognition 10.64898/2026.05.05.722899 medRxiv
Top 0.1%
1.5%
Show abstract

In contrast to static formalisms, computational definitions describe the operational mechanisms of a model. Simulations are an essential part of the cycle of theory development and refinement, assisting researchers in formulating the precise definitions that models require, and making accurate predictions. This manuscript introduces a computational implementation of Pavlovian learning models in a Python environment, termed Pavlovian Associative Learning Models Simulation (PALMS). In addition to the canonical Rescorla-Wagner model, attentional approaches are implemented, including Pearce-Kaye-Hall, Mackintosh Extended, Le Pelleys Hybrid, and a novel extension of the Rescorla-Wagner model featuring a unified variable learning rate that synthesises Mackintoshs and Pearce and Halls opposing conceptualisations. To our knowledge, only the first attentional model has been previously specified computationally in a general design tool. PALMS integrates a graphical interface that permits the input of entire experimental designs in an alphanumeric format, akin to that used by experimental neuroscientists. It uniquely enables the simulation of experiments involving hundreds of stimuli, such as those used with human participants, and the computation of configural cues and configural-cue compounds across all models, thereby substantially broadening their predictive capabilities. A comprehensive description of the models implementation and the environment functionalities is provided in the paper; these include efficient and accurate operation and instant visualisation of predicted results across different models within a single architecture and environment. We evaluate PALMS by simulating five published experiments in the associative learning literature that assessed the predictive scope of existing models, and we show that this implementation provides neuroscientists with a useful tool for identifying critical variables, refining experimental designs, making precise predictions, comparing model fitness, and formulating new theoretical approaches. PALMS is licensed under the open-source GNU Lesser General Public License 3.0. The environment source code and the latest multiplatform release build are accessible as a GitHub repository at https://github.com/cal-r/PALMS-Simulator. Author summaryResearch on associative learning is multidisciplinary, encompassing disciplines such as neuroscience, AI, psychology, psychiatry, behavioural sciences, planning, and marketing. Unlike static formalisms, precise computational definitions specify how a model operates, enabling model simulation, swift and error-free prediction calculations, which are essential for testing theories, comparing predictions, holding models accountable, and providing a common language across fields. We introduce Pavlovian Associative Learning Models Simulation (PALMS), a user-friendly, open-source Python environment for simulating classical conditioning and studying the role of attention in learning. PALMS implements the prescriptive Rescorla-Wagner and attentional models: Pearce-Kaye-Hall, Mackintosh Extended, Le Pelleys Hybrid, and a new hybrid model with a unified variable learning rate that blends Mackintosh and Pearce-Halls conflicting views. Its graphical interface makes it easy for neuroscientists to enter experiments. Our computational implementation supports simulations with hundreds of stimuli, configural cues, and compounds, broadening the models predictive power. Designed for efficiency, it offers instant visual results and useful features. We evaluate PALMS by simulating five published experiments, highlighting its value for model comparison and refinement, and, more generally, as a tool to assist research.

13
pykarambola: Minkowski tensor morphometry of 3D structures

Khurana, Y.; Ishihara, K.

2026-06-18 bioinformatics 10.64898/2026.06.16.730752 medRxiv
Top 0.1%
1.5%
Show abstract

Three-dimensional biological morphologies encode functional and physiological state, yet the directional, orientational, and topological properties of these shapes are rarely captured by morphometric tools available for bioimage analysis. Minkowski tensors are mathematically rigorous tensor-valued measures that encode surface curvature and directionality for objects of arbitrary topology, with tensor eigensystems that directly quantify elongation axes and anisotropy. A C++ implementation, karambola (1), computes Minkowski tensors for triangulated surfaces but is inaccessible within Python-based bioimage workflows. Here we present pykarambola, a pip installable Python package that accepts NumPy arrays and standard mesh formats and returns Minkowski tensors, including derived anisotropy and orientation quantities. A high-level label-image API converts 3D integer arrays into per-object Minkowski tensors in a single call, making pykarambola directly compatible with the output of widely used segmentation tools. An optional Cython extension accelerates graph-traversal steps of mesh initialization for large-scale analyses. Benchmarked on 1,584 adrenal gland meshes, pykarambola reproduces all 121 C++ karambola output features to near-floating-point agreement and, in the pure-Python build, is 2.8x faster at 283 and 1.5x faster at 643 voxel resolution, with speedups primarily attributable to karambolas sequential per-object file I/O. pykarambola is freely available as an open-source software package.

14
DigitalPedon: A Novel Digital Twin Framework for Soil Profile Monitoring and Global Soil Data Interoperability

Youssef, A.; Badreldin, N.

2026-05-08 bioengineering 10.64898/2026.05.05.722891 medRxiv
Top 0.1%
1.5%
Show abstract

The Digital Pedon (DP) is an open-source Python framework that represents a soil profile as a continuously updated digital twin, bridging three persistent gaps in soil science: disconnected models and observations, cross-database interoperability, and the inference gap between raw sensor signals and agronomically meaningful variables. Integrating real-time sensor streams, model-based solver chains (Model-Zoo), GLOSIS-compliant ontology mapping, and a novel LLM agentic interface layer enabling natural language soil queries, the DP supports applications spanning precision agriculture, digital soil mapping, and environmental sustainability assessment. Four proof-of-concept experiments confirm automatic profile initialisation fidelity, solver chain consistency, ontology compliance, and user-defined solver extensibility.

15
pylimma: a faithful, AnnData-native Python port of R limma for differential expression analysis

Mulvey, J.

2026-07-10 bioinformatics 10.64898/2026.07.06.736732 medRxiv
Top 0.1%
1.5%
Show abstract

pylimma is a faithful Python port of limma, intended to bring one of the most widely used tools for differential expression analysis to the developing Python ecosystem for transcriptomics and proteomics. We validated pylimma against the existing R implementation through 227 function-level comparisons and across six real world datasets spanning microarray, RNAseq, proteomics and single-cell transcriptomics. pylimma reproduces limmas numerical output to a median agreement of 13 significant figures and calls identical sets of differentially expressed features and gene sets. This supports its use as a drop-in replacement for the R implementation.

16
Figra: A WebAssembly-based Excel Add-in for publication-quality scientific visualization with ggplot2

Sato, Y.

2026-05-12 bioinformatics 10.64898/2026.05.06.723320 medRxiv
Top 0.1%
1.4%
Show abstract

Data visualization is a critical step in scientific communication. Most researchers rely on subscription-based software for this purpose, which requires ongoing licensing costs. Free alternatives such as R and Python offer publication-quality output but demand programming expertise that many researchers do not possess. Artificial intelligence tools can assist with figure generation but remain frustrating when users wish to fine-tune specific visual parameters to their preference. Meanwhile, Microsoft Excel, the most widely used tool for scientific data storage and management, offers limited visualization capabilities, forcing researchers to transfer their data to external software as an extra step before creating figures. Here we present Figra, a free Excel Office Add-in that eliminates this extra step by enabling publication-quality ggplot2-based figure generation directly within Excel, with simple and direct control over every visual option. Figra leverages WebAssembly technology (webR) to execute R code entirely within the browser, requiring no R installation, no subscription, and no server connection. The add-in supports over 20 chart types spanning distribution plots, grouped comparisons, time-series, scatter plots, and specialized curve-fitting analyses. For applicable chart types, Figra performs automated or manual statistical analysis supporting both paired and unpaired designs across two or more groups. Additionally, Figra exports simplified, executable R code that reproduces the displayed figure, serving as an educational tool for researchers wishing to learn ggplot2. Figra is open-source and freely available at https://h20gg702.github.io/figra-pages/index.html while the source code is provided at https://github.com/h20gg702/Figra.

17
cran2crux: automatically create CRUX ports for R-packages

Petrov, P.; Izzi, V.

2026-05-13 bioinformatics 10.64898/2026.05.09.723963 medRxiv
Top 0.2%
1.3%
Show abstract

MotivationR together with CRAN and Bioconductor provides one of the richest ecosystems for bioinformatics and computational biology, with thousands of specialized packages. While GNU/Linux is a vastly-used operating system in this field, R-packages are typically managed independently of the systems native package manager. This separation makes installation, updates and mass rebuilds cumbersome. CRUX, a minimalist semi-source GNU/Linux distribution, offers great flexibility with its ports-based system for the seamless integration of R-packages with its native package manager. ResultsThe hereby presented cran2crux tool automatically generates CRUX ports for packages from both CRAN and Bioconductor. It performs recursive dependency resolution, handles naming conventions, extracts dependencies information, and supports inclusion of optional dependencies. The tool also provides convenient functions for checking updates and regenerating outdated ports. It can generate over 140 ports for complex packages such as Seurat in approximately 11 seconds, dramatically simplifying the maintenance of large R-dedicated repositories on CRUX. Availabilitycran2crux is available under the MIT license at https://github.com/izzilab/cran2crux. As of now, more than 650 R package ports, generated with the tool, are available in the CRUX ports database.

18
OncoContour: An Interactive Platform for Geographic Visualization and Demographic Analysis of Cancer Incidence.

White, D.; Uzun, A.

2026-05-22 bioinformatics 10.64898/2026.05.20.726625 medRxiv
Top 0.2%
1.2%
Show abstract

Cancer incidence varies substantially across geographic regions and demographic groups, yet translating large-scale surveillance datasets into accessible, interpretable visualizations remains a challenge for researchers and public health professionals without computational expertise. We developed OncoContour, an interactive web-based platform that enables geographic visualization and demographic analysis of cancer incidence data through a browser-accessible interface. To demonstrate its capabilities, we analyzed publicly available cancer incidence data from the United States Cancer Statistics database via CDC WONDER, covering five major cancer types across four northeastern U.S. metropolitan statistical areas from 2017 through 2022, supplemented by demographic data from the U.S. Census Bureau American Community Survey. OncoContour integrates population distribution heatmaps, per-capita cancer incidence heatmaps, interactive multi-city temporal trend charts, structured cancer data tables, and demographic visualizations covering race, ethnicity, age, and sex distributions into a single dynamically generated HTML report. The platform is implemented in Python using Flask, Folium, Plotly, and Matplotlib, and is containerized using Docker for reproducible local deployment. Across all four metropolitan areas, breast and prostate cancers accounted for the highest incidence counts over the study period, while a decline in reported cases observed in 2020 is consistent with documented disruptions to cancer screening during the COVID-19 pandemic. By integrating geospatial mapping, temporal analysis, and demographic visualization within a unified, no-code interface, OncoContour aims to support cancer surveillance, epidemiological investigation, and targeted public health planning. OncoContour is freely available at https://github.com/alperuzun/oncocontour_docker.

19
3dcon: tomogram denoising by deconvolution

Kirchweger, P.; Melnikovsky, L.; Seifer, S.; Elbaum, M.

2026-06-18 biochemistry 10.64898/2026.06.15.732138 medRxiv
Top 0.2%
1.1%
Show abstract

Cryo-electron tomography is an expanding technology for the study of macromolecules, viruses, and cells. It is often applied to specimens that are too large or heterogeneous for methods based on 2D image averaging such as single particle analysis, e.g., intracellular membranes or organelles. Current practice records a tilt series of projection images in rotation. Reconstruction is normally an ill-posed mathematical problem. Particularly for the under-determined case of sparse data, discrete tilt angles, and a limited tilt range, characteristic artifacts appear in the reconstructed slices. Much of what appears as noise is in fact structural: the projection of contrast from different planes. Various schemes are employed to regularize the reconstruction, including machine-learning frameworks built on neural networks. To the extent that the noise is structural, it might be suppressed by deconvolution with a suitable kernel. This was demonstrated and has been used regularly in cryo-STEM tomography of thick specimens where the under-sampling problem is particularly acute. Here we present 3dcon as an open-source extension of the entropy-regularized deconvolution algorithm that had been adopted from fluorescence microscopy. It takes advantage of modern computing hardware for convenient and fast processing. Deconvolution is entirely algorithmic, meaning that successful processing of the data does not depend on the data itself. As such it should be robust in a wide variety of applications.

20
DAQplugin: Deep Learning based Real-time Model Evaluation Plugin for ChimeraX

Terashi, G.; Zhu, H.; Kihara, D.

2026-06-15 bioinformatics 10.64898/2026.06.11.731735 medRxiv
Top 0.2%
1.1%
Show abstract

Although an increasing number of protein structures are determined by cryogenic electron microscopy (cryo-EM), protein structure modeling frequently suffers from residue misassignments and sequence register shifts, particularly in regions with ambiguous density. Here, we present DAQplugin, a ChimeraX plugin that performs real-time evaluation of protein models against cryo-EM density maps using the deep-learning-based residue-wise model quality (DAQ) score. Unlike existing validation tools that are typically applied after model construction, DAQplugin enables real-time deep-learning-based validation during model building and refinement. To our knowledge, DAQplugin is the first tool that provides real-time deep-learning based validation of protein models for cryo-EM map within an interactive modeling environment. In addition to identifying potential modeling errors, DAQplugin also provides guidance for correcting sequence register shifts by suggesting alternative residue placements along the backbone. The computation in this plugin is designed to run efficiently on general CPUs without requiring GPU hardware. Using DAQplugin, users can perform deep-learning-based validation on standard laptops during interactive model building, model-map fitting, and refinement. DAQplugin is able to facilitate more accurate interpretation of cryo-EM density maps and improve the reliability assessment of protein structure models. SynopsisDAQplugin provides real-time residue-wise validation of protein models with cryo-EM maps in ChimeraX.