Back

Nature Machine Intelligence

Springer Science and Business Media LLC

All preprints, ranked by how well they match Nature Machine Intelligence's content profile, based on 70 papers previously published here. The average preprint has a 0.10% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Grounding olfactory perception in language: Benchmarks and models for generating natural language odor descriptions

Mascart, C.; Tran, K.; Samoilova, K.; Storan, L. T.; Liu, T.; Koulakov, A.

2026-03-05 animal behavior and cognition 10.64898/2026.03.04.709650 medRxiv
Top 0.1%
40.1%
Show abstract

Recent advances in deep learning have enabled prediction of odorant perception from molecular structure, opening new avenues for odor classification. However, most existing models are limited to predicting percepts from fixed vocabularies and fail to capture the full richness of olfactory experience. Progress is further limited by the scarcity of large-scale olfactory datasets and the lack of standardized metrics for evaluating free-form natural-language odor descriptions. To address these challenges, we introduce Odor Description and Inference Evaluation Understudy (ODIEU), a benchmark which includes perceptual descriptions of over 10,000 molecules paired with a model-based metric for evaluating free-form odor text descriptions. The model-based metric uses Sentence-BERT (SBERT) models which are fine-tuned on olfactory descriptions to allow better evaluation of human-generated odor texts. Using the fine-tuned SBERT models, we show that free-form text odor descriptions contain additional perceptual information in their syntactic structure compared to semantic labels. We further introduce CIRANO (Chemical Information Recognition and Annotation Network for Odors), a transformer-based model that generates free-form odor descriptions directly from molecular structure, thus implementing the molecular structure-to-text (S2T) prediction. CIRANO achieves performance comparable to humans. Finally, we generate human-like descriptions from mouse olfactory bulb neural data using an invertible SBERT model, yielding neural-to-text (N2T) predictions highly aligned with human descriptions. Together, CIRANO and ODIEU establish a standardized framework for generating natural language olfactory descriptions and evaluating their alignment with human perception. Code is available at https://github.com/KoulakovLab/ODIEU

2
DamageFormer: a damage-aware multimodal deep learning framework for DNA lesion identification from nanopore sequencing

Yang, Q.; Li, L.; Ma, Q.; Yin, R.

2026-05-18 genomics 10.64898/2026.05.14.725245 medRxiv
Top 0.1%
39.3%
Show abstract

BackgroundDNA lesions arise from endogenous metabolism and environmental exposure and are the major drivers of mutagenesis, aging, and cancer development. However, mapping DNA damage at nucleotide resolution remains a technically challenging task. Nanopore sequencing enables direct detection of chemical perturbations through alterations in ionic current signals. Despite this potential, existing computational approaches remain limited in their capacity to generalize across diverse lesion types and to effectively integrate nucleotide sequence context with raw signal information for accurate detection and localization. ResultsWe presented DamageFormer, a multimodal deep learning framework for detection and localization of DNA lesions using native nanopore sequencing data. Central to this framework is LesionBERT, a damage-aware genomic foundation model built upon DNABERT-2 and enhanced with lesion-focused reconstruction objectives to improve representation of chemically modified bases. DamageFormer integrated LesionBERT with a neural signal model through an adaptive gating mechanism, enabling dynamic weighting of sequence context and nanopore signal evidence. The model was trained using a joint objective that combines prediction, localization, and contrastive alignment losses to promote cross-modal coherence and spatial precision. On an oxidative DNA damage benchmark comprising paired sequence and signal data, DamageFormer achieved an AUROC of 0.99997 for lesion detection and a mean absolute localization error of 0.00439, consistently outperforming state-of-the-art baselines. Model interpretation analyses revealed context-dependent modality weighting that adapts to variation in signal quality and sequence ambiguity. The proposed framework further generalizes to chemically distinct guanine lesions not observed during the training process, demonstrating its robustness and transferability to unseen damage types. ConclusionsDamage-aware biological language modeling combined with adaptive multimodal fusion enables accurate and interpretable identification of DNA lesions from nanopore sequencing data. This framework provides a scalable approach for characterizing genome-wide damage landscapes and illustrates how chemical DNA information can be systematically incorporated into genomic language models. The source code and pretrained models of this work are available at: https://github.com/UF-HOBIYin-Lab/DamageFormer.

3
Training-free Design of Deep Networks as Ensembles of Clinical Experts

Wu, T.; Chen, W.; Zhang, Z.

2024-03-18 health informatics 10.1101/2024.03.17.24304438 medRxiv
Top 0.1%
38.5%
Show abstract

Artificial intelligence (AI) techniques such as deep learning hold tremendous potential for improving clinical practice. However, clinical data complexity and the need for extensive specialized knowledge represent major challenges in the current, human-driven model design. Moreover, as human interpretation of a clinical problem is inherently encoded in the model, the conventional single model paradigm is subjective and cannot fully capture the prediction uncertainty. Here, we present a fast and accurate framework for automated clinical deep learning, TEACUP (training-free assembly as clinical uncertainty predictor). The core of TEACUP is a newly developed metric that faithfully characterizes the quality of deep networks without incurring any cost for training of these networks. When compared to conventional, training-based approaches, TEACUP reduces computation costs by more than 50% while achieving improved performance across distinct clinical tasks. This efficiency allows TEACUP to create ensembles of expert AI models, contributing to recommendations in clinical practice by mimicking the approach of using multiple human experts when interpreting medical data. By combining multiple perspectives, TEACUP provides more robust predictions and uncertainty quantification, paving the way for more reliable clinical AI.

4
SC-MAMBA2: Leveraging State-Space Models for Efficient Single-Cell Ultra-Long Transcriptome Modeling

Zhao, Y.; Zhao, B.; Zhang, F.; He, C.; Wu, W.; Lai, L.

2024-10-26 cell biology 10.1101/2024.09.30.615775 medRxiv
Top 0.1%
34.4%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWThe rapid advancement of single-cell sequencing technology has significantly deepened our understanding of cellular heterogeneity, yet it concurrently presents substantial challenges for the unified modeling of single-cell data. Simultaneously, pre-trained foundation models have achieved notable success in domains such as natural language processing and image analysis. However, extending these models to accommodate ultra-long single-cell transcriptome sequences, characterized by an extensive number of genes, remains a formidable task. In this study, we introduce SC-MAMBA2, based on the MAMBA2 architecture, meticulously designed with a bidirectional modeling approach tailored for single-cell transcriptomics data. As the first single-cell foundation model to integrate state-space models (SSMs) underlying MAMBA2 architecture, SC-MAMBA2 features over 625 million parameters, covers more than 60,000 genes, and was pre-trained on a dataset of over 57 million cells, making it the most comprehensive solution for processing ultra-long transcriptome sequences. Extensive bench-marking across a diverse array of downstream tasks consistently demonstrates that SC-MAMBA2 surpasses state-of-the-art models, delivering superior accuracy and enhanced computational efficiency.

5
ProtJEPA: A Multimodal Joint-Embedding Predictive Architecture for Protein Biological World Modeling with Multi-TeacherModality-Attentive Fusion

Ravideshik, V. L.; Kim, J.; Kellis, M.

2026-08-09 cell biology 10.64898/2026.08.03.742606 medRxiv
Top 0.1%
34.4%
Show abstract

Over 99.9% of known protein sequences lack experimentally validated functional annotations. We present ProtJEPA, a multimodal Joint-Embedding Predictive Architecture that trains a sequence-only student encoder to predict joint embeddings spanning ten biological modalities--sequence, structure, knowledge graph, protein interactions, literature, localization, tissue expression, GO function, anatomy, and disorder--requiring only sequence at inference. The key innovation is target whitening, which eliminates severe anisotropy in joint targets (mean cosine 0.984 to 0.086) and prevents representation collapse without covariance regularization. On 1,828 held-out dark proteins with zero primary Pfam family overlap with training, ProtJEPA achieves 58.07% Hit@10 on zero-shot GO retrieval (+2.80 pp, p = 0.020), 69.99% enzyme class accuracy (+9.64 pp, p < 0.001), and +11.87 pp subcellular localization at 1% labels (p < 0.001). Under realistic dark-protein deployment conditions where relational modalities are unavailable, ProtJEPA significantly outperforms naive concatenation of remaining modalities. Cross-domain evaluations on drug-target interaction and disorder prediction confirm transfer beyond training modalities, with the T1-only < ESMC < ProtJEPA ordering replicated across six independent tasks. Ablations establish that Phase 1 aggregator pretraining and target whitening are each independently load-bearing.

6
TabSyM: A Generative Pipeline for Small Multi-Cohort Omics Tabular Data

Yu, N.; Wang, Y.; Olsen, L. K.; Zhang, B.; Zhang, h.; Liu, Z.

2025-07-18 bioinformatics 10.1101/2025.07.14.664738 medRxiv
Top 0.1%
34.3%
Show abstract

Machine learning applications in biomedicine such as omics data analysis are frequently hindered by datasets that are small, high-dimensional, and affected by batch effects across different patient cohorts. To address these challenges, we introduce TabSyM, a modular generative pipeline that synthesizes high-quality, task-relevant data to improve predictive modeling. TabSyM integrates three key stages: it extends a diffusion-based model (TabDDPM) to generate new omics data, employs a novel task-aware sampling mechanism guided by Bayesian optimization to select the most informative synthetic samples, and uses a Multi-Domain Adversarial Network (MDAN) to align data distributions for cross-cohort generalization. We validated our pipeline on a challenging, real-world task of predicting 3-year survival in gastric cancer patients from high-dimensional scRNA-seq data across five cohorts. The full TabSyM pipeline achieved a 30.2% AUROC improvement over the best tree-based models and an 11.5% AUROC gain over leading automated machine learning frameworks. Furthermore, the generative and sampling components are model-agnostic and can substantially boost the performance of classical models like XGBoost independently. These results establish that combining generative modeling with task-aware sampling and domain adaptation provides a robust and effective strategy for overcoming critical data limitations in biomedical tabular data analysis.

7
GraphTox: A Semi-Supervised Pre-Trained Framework for Peptide Toxicity Prediction using Geometric Graph Transformer and LORA-Based Finetuning

BHADURI, S.; Das, D.; MITRA, P.

2026-05-27 bioinformatics 10.64898/2026.05.23.727225 medRxiv
Top 0.1%
34.3%
Show abstract

Peptides are widely used as potential therapeutic agents in drug discovery and biotechnology because they are specific, effective, and relatively inexpensive to produce. They are used in drug development, vaccines, and antimicrobial treatments. However, peptide toxicity remains a major concern as it offers unwanted toxic consequences, such as membrane rupture, haemolysis, tissue damage and adverse immunological response. Early detection of toxic peptide candidates is vital for the development of safe and effective therapies. Current computational methods for predicting peptide toxicity are largely based on hand-crafted sequence descriptors or sequence-only deep learning architectures that may not fully account for the underlying 3-dimensional structural determinants of peptide toxicity. We introduce GraphTox, a structure-aware geometric deep learning framework which combines self-supervised graph representation learning with hierarchical structural modelling to accurately predict peptide toxicity. Our framework learns geometry-aware embeddings from peptide structural graphs via self-supervised masked residue reconstruction, based on a Masked Graph Autoencoder (MGAE) built on a Geometric Graph Transformer (GGT) encoder. The pretrained structural representations are cross fused via a multi-scale U-Net architecture to capture both local residue-level interactions and global conformational patterns associated with peptide toxicity. GraphTox explicitly models spatial relationships between residues, thereby efficiently capturing structural aspects that are generally neglected by sequence-based predictors, such as residue clustering, hydrophobic interactions and electrostatic organization. On benchmark datasets our framework shows superior performance and interpretability over the existing state-of-the-art methods. Our hybrid hierarchical structural modelling framework is a superior computational platform to improve the prediction of peptide toxicity and expedite the creation of safer peptide therapies. https://github.com/debraj-55555/GraphTox

8
GEM-GPT Enables Personalized Cell Type-Resolved Therapeutic Design for Systems Pharmacology

Zhang, S.; Ohlan, R.; Mottaqi, M.; Xie, L.

2026-07-22 pharmacology and toxicology 10.64898/2026.07.17.739269 medRxiv
Top 0.1%
32.9%
Show abstract

Generative artificial intelligence (AI) has emerged as a powerful framework for drug discovery, yet most current approaches follow one-drug-one-gene target-based paradigms that struggle to capture the complexity and heterogeneity of chronic and systemic diseases. Omics-driven systems pharmacology provides a promising strategy to overcome these limitations, but generative AI tools specifically designed for systems pharmacology-oriented drug design remain scarce. To address this gap, we introduce GEM-GPT, a transcriptomics-based molecule generation framework that designs personalized therapeutic compounds capable of reverting cell type-specific disease states back to a healthy phenotype. GEM-GPT employs a biology-inspired deep fusion architecture that couples a single-cell RNA-sequencing (scRNA-seq) foundation model with a molecular GPT model, enabling the modeling of cell type-specific chemical-gene interactions during molecule generation. This integration allows GEM-GPT to outperform state-of-the-art baselines, generate distinct molecules for different cell types, and generalize robustly to previously unseen cell types. We further demonstrate the utility of GEM-GPT through a case study in personalized drug discovery for opioid use disorder (OUD). In this application, GEM-GPT successfully identifies both therapeutic compounds possessing distinct chemotypes and existing FDA-approved drugs predicted to modulate cell type-specific OUD disease phenotypes in individual patients. Together, these results establish GEM-GPT as an advance in AI-driven systems pharmacology by bridging single-cell omics and molecular generation to support personalized, systems-aware therapeutic design.

9
Sparse autoencoder features from InterPLM predict neuropeptide precursors among secreted proteins

Kulikova, A. V.; Bookout, A. L.; Koch, T. L.; Safavi-Hemami, H.

2026-08-21 bioinformatics 10.64898/2026.08.20.746077 medRxiv
Top 0.1%
32.4%
Show abstract

Neuropeptides are a diverse class of short, secreted signaling molecules that regulate key physiological processes in animals. Despite their important biological roles and increasingly recognized therapeutic value, the discovery of new neuropeptides remains challenging, largely because their short length and high sequence heterogeneity limit the effectiveness of motif- and homology-based approaches. Here, we present a pipeline for neuropeptide precursor prediction that leverages sparse autoencoders (SAEs) from the protein language model InterPLM to decode dense protein language model embeddings into sparse, disentangled features. We identify a small subset of features strongly associated with neuropeptide precursors that achieve high discriminative performance. A logistic regression classifier trained on this reduced feature set, accurately separates human neuropeptide and non-neuropeptide sequences. We then applied this classifier to important model organisms: mouse (Mus musculus), zebrafish (Danio rerio), nematode (Caenorhabditis elegans), and fruit fly (Drosophila melanogaster ) and show that the approach generalizes across diverse species. Overall, InterPLM SAE features provide an interpretable and effective strategy for neuropeptide prediction and enable a trained classifier to predict neuropeptides from large datasets. A web tool for this classifier is freely available at https://biolib.com/ATGCACTGTTCAGGCCTC/SAE-Neuropeptide-Predictor

10
RulePep: Interpretable ESM-Guided Neural-Symbolic Peptide Classification

Midjani, F.; Ghelich, R.; Keshtkar, F. Z.; Malekpour, M.; Lee, H.

2026-07-06 bioinformatics 10.64898/2026.07.03.736448 medRxiv
Top 0.1%
31.3%
Show abstract

Peptides are increasingly explored as therapeutic candidates, delivery vectors, and functional biomolecules, but experimental screening of peptide activity and safety remains costly because the sequence space is vast and small sequence changes can alter functionality. Computational peptide classification can therefore help prioritize candidates. However, many protein-language-model-based classifiers achieve strong performance using opaque prediction heads, making it difficult to determine which learned evidence supports or opposes a prediction. We present RulePep, an ESM-2-guided neural-symbolic classifier for peptide-function prediction. RulePep maps frozen ESM-2 sequence representation to learned latent predicates, polarity-constrained differentiable rules, and an additive symbolic logit whose components can be inspected at the case level. We evaluate RulePep on three biologically distinct peptide classification tasks: blood-brain barrier penetration, hemolytic potency, and anticancer activity. On the BBPpredict, HemoPI3, and AntiCP 2.0 alternate benchmark datasets, RulePep achieved AUROC/MCC values of 0.8869/0.6850, 0.9155/0.6820, and 0.9765/0.8633, respectively. Ablation experiments supported the contributions of multi-layer representation pooling, rule polarity, mined-rule initialization, symbolic capacity, and rule-derived aggregation. RulePep combines competitive predictive performance with additive logit reconstruction, rule-level evidence reporting, and predicate-suppression auditing, providing a transparent sequence-based framework for peptide candidate prioritization.

11
Sequence-Based Therapeutic Peptide Classification with Augmented Negative Sampling

Ellerbrock, R.; Valentini, A.; Paul, A. C.; Mukhopadhyay, S.; Perelshtein, M. R.

2026-06-11 bioinformatics 10.64898/2026.06.07.730473 medRxiv
Top 0.1%
31.2%
Show abstract

Therapeutic peptides offer high target specificity, low toxicity, and the ability to modulate protein-protein interactions, yet experimental functional characterization remains costly and slow. Computational prediction of therapeutic function directly from sequence could accelerate peptide screening and enable generative design pipelines, but requires reliable discrimination between therapeutic and non-therapeutic peptides. Existing multi-label predictors cover few functions, rely on limited datasets, and exhibit high False Positive Rates (FPRs), limiting their practical utility. We present a lightweight CNN classifier trained on the most comprehensive therapeutic peptide database to date (54,655 peptides, 48 functional categories). A key contribution is a statistically motivated negative sampling strategy using Markov models to generate diverse synthetic decoys at multiple difficulty levels. When evaluated on this controlled decoy benchmark, the FPR is reduced from over 60% for previous models to 2.1% for our approach. On positive therapeutic samples, our fine-tuned five-model ensemble achieves 79.9% Micro F1 and 54.6% Macro F1 while requiring only amino acid sequences as inputs. Analysis using a sparse L1-constrained variant of our model shows that convolutional filters capture conserved functional motifs and statistically improbable non-therapeutic patterns, with downstream layers combining these signals, providing mechanistic evidence that the network learns biologically meaningful structure. On an external generalization benchmark derived from TPpred-LE, our model achieves 55.3% Micro F1 and 38.6% Macro F1 on the 12 shared labels, close to the benchmark-specific baseline (57.9%/38.1%), while retaining substantially broader therapeutic label coverage. Code and models will be made available at https://github.com/terra-quantum-public/tq-therapep-ai.

12
Biological network-inspired interpretable variational autoencoder

Seninge, L.; Anastopoulos, I.; Ding, H.; Stuart, J.

2020-12-19 bioinformatics 10.1101/2020.12.17.423310 medRxiv
Top 0.1%
31.1%
Show abstract

Deep learning architectures such as variational autoencoders have revolutionized the analysis of transcriptomics data. However, the latent space of these variational autoencoders offers little to no interpretability. To provide further biological insights, we introduce a novel sparse Variational Autoencoder architecture, VEGA (Vae Enhanced by Gene Annotations), whose decoder wiring is inspired by a priori characterized biological abstractions, providing direct interpretability to the latent variables. We demonstrate the interpretability and flexibility of VEGA in diverse biological contexts, by integrating various sources of biological abstractions such as pathways, gene regulatory networks and cell type identities in the latent space of our model. We show that our model could recapitulate the mechanism of cellular-specific response to treatments, the status of master regulators as well as jointly investigate the cell type and cellular state identity in developing cells. We envision the approach could serve as an explanatory biological model in contexts such as development and drug treatment experiments.

13
NeuroNarrator: A Generalist EEG-to-Text Foundation Model for Clinical Interpretation via Spectro-Spatial Grounding and Temporal State-Space Reasoning

Wang, G.; Yang, S.; Ding, J.-e.; Zhu, H.; Liu, F.

2026-03-10 bioinformatics 10.64898/2026.03.07.707799 medRxiv
Top 0.1%
31.1%
Show abstract

Electroencephalography (EEG) provides a non-invasive window into neural dynamics at high temporal resolution and plays a pivotal role in clinical neuroscience research. Despite this potential, prevailing computational approaches to EEG analysis remain largely confined to task-specific classification objectives or coarse-grained pattern recognition, offering limited support for clinically meaningful interpretation. To address these limitations, we introduce NeuroNarrator, the first generalist EEG-to-text foundation model designed to translate electrophysiological segments into precise clinical narratives. A cornerstone of this framework is the curation of NeuroCorpus-160K, the first harmonized largescale resource pairing over 160,000 EEG segments with structured, clinically grounded natural-language descriptions. Our architecture first aligns temporal EEG waveforms with spatial topographic maps via a rigorous contrastive objective, establishing spectro-spatially grounded representations. Building on this grounding, we condition a Large Language Model through a state-space-inspired formulation that integrates historical temporal and spectral context to support coherent clinical narrative generation. This approach establishes a principled bridge between continuous signal dynamics and discrete clinical language, enabling interpretable narrative generation that facilitates expert interpretation and supports clinical reporting workflows. Extensive evaluations across diverse benchmarks and zero-shot transfer tasks highlight NeuroNarrators capacity to integrate temporal, spectral, and spatial dynamics, positioning it as a foundational framework for time-frequency-aware, open-ended clinical interpretation of electrophysiological data.

14
GeoEPred: A Multimodal Structure-Aware Geometric Deep Learning Framework for Gram-Negative Bacterial Secreted Effector Prediction with Sequence Semantics

Song, S.; Shi, H.; Wu, H.; Liu, D.; Lin, Y.; Mat Isa, N. A.; Zou, Q.; Wei, L.

2026-05-20 genomics 10.64898/2026.05.18.725929 medRxiv
Top 0.1%
31.0%
Show abstract

Accurate prediction of effector proteins secreted by Gram-negative bacteria is important for elucidating bacterial pathogenic mechanisms and developing precise anti-infective strategies. Although existing methods have benefited from the strong sequence feature extraction capacity of pretrained protein language models, reliance on linear sequence information alone often fails to fully capture the three-dimensional conformational signals required for virulence functions. Meanwhile, conventional structure-based methods are limited by the scarcity of experimentally resolved protein structures. To address these challenges, We propose GeoEPred, a multimodal deep learning framework designed for the synergistic modeling of protein sequence and structure to identify Gram-negative bacterial effector proteins. Specifically, the model integrates sequence-contextual embeddings from a pretrained protein language model with three-dimensional structural representations predicted by ESMFold. A feature projection network refines fine-grained sequence signals associated with effector functions, while geometric vector perceptrons characterize inter-residue orientations, distances, and local spatial topology to capture potential structural conformational motifs. To further enable effective cross-modal fusion, we design a cross-modal alignment and feature-tokenized self-attention module. This module enhances consistency between the sequence-semantic and structural-geometric spaces through contrastive learning and models associations between linear functional motifs and spatial conformational patterns at a fine-grained token level. Extensive evaluations on multiple benchmark datasets show that GeoEPred achieves better predictive performance than existing leading models in T3SE, T4SE, and T6SE prediction tasks, while maintaining stable performance in remote homolog recognition scenarios. Moreover, the modular and extensible architecture of GeoEPred demonstrates strong generalization ability and substantial application potential for genome-scale effector protein discovery. Author summarySecreted effector proteins are central virulence factors used by many Gram-negative bacterial pathogens to execute infection strategies. Their functions are governed not only by secretion signals and short linear motifs in the amino acid sequence, but also by three-dimensional folds, local domains, and surface geometric patterns. However, current predictors mainly exploit sequence-contextual features, limiting their ability to model the correspondence between linear sequence signals and spatial conformational motifs, and thereby constraining accuracy and interpretability. Here, we present GeoEPred, a multimodal deep learning framework for secreted effector protein identification. GeoEPred couples sequence-semantic embeddings from a pretrained protein language model with structural representations learned by geometric vector perceptrons. A cross-modal alignment and interaction module uses contrastive learning to improve functional consistency between sequence and structure modalities, while feature-token attention captures fine-grained links between key linear and conformational motifs. Across benchmark datasets covering multiple effector types, GeoEPred outperforms existing state-of-the-art methods and provides interpretable evidence from sequence fragments, structural regions, and cross-modal associations, supporting functional annotation, pathogenic mechanism analysis, and experimental validation.

15
Flexible and Highly-Efficient Feature Perception for Molecular Traits Prediction via Self-interactive Deep Learning

Hu, Y.; Sirinukunwattana, K.; Li, B.; Gaitskell, K.; Bonnaffe, W.; Wojciechowska, M.; Wood, R.; Alham, N. K.; Malacrino, S.; Woodcock, D.; Verrill, C.; Ahmed, A.; Rittscher, J.

2023-08-05 pathology 10.1101/2023.07.30.23293391 medRxiv
Top 0.1%
30.7%
Show abstract

Predicting disease-related molecular traits from histomorphology brings great opportunities for precision medicine. Despite the rich information present in histopathological images, extracting fine-grained molecular features from standard whole slide images (WSI) is non-trivial. The task is further complicated by the lack of annotations for subtyping and contextual histomorphological features that might span multiple scales. This work proposes a novel multiple-instance learning (MIL) framework capable of WSI-based cancer morpho-molecular subtyping across scales. Our method, debuting as Inter-MIL, follows a weakly-supervised scheme. It enables the training of the patch-level encoder for WSI in a task-aware optimisation procedure, a step normally improbable in most existing MIL-based WSI analysis frameworks. We demonstrate that optimising the patch-level encoder is crucial to achieving high-quality fine-grained and tissue-level subtyping results and offers a significant improvement over task-agnostic encoders. Our approach deploys a pseudo-label propagation strategy to update the patch encoder iteratively, allowing discriminative subtype features to be learned. This mechanism also empowers extracting fine-grained attention within image tiles (the small patches), a task largely ignored in most existing weakly supervised-based frameworks. With Inter-MIL, we carried out four challenging cancer molecular subtyping tasks in the context of ovarian, colorectal, lung, and breast cancer. Extensive evaluation results show that Inter-MIL is a robust framework for cancer morpho-molecular subtyping with superior performance compared to several recently proposed methods, even in data-limited scenarios where the number of available training slides is less than 100. The iterative optimisation mechanism of Inter-MIL significantly improves the quality of the image features learned by the patch embedded and generally directs the attention map to areas that better align with experts interpretation, leading to the identification of more reliable histopathology biomarkers.

16
Robust Neural Networks are More Interpretable for Genomics

Koo, P. K.; Qian, S.; Kaplun, G.; Volf, V.; Kalimeris, D.

2019-06-03 genomics 10.1101/657437 medRxiv
Top 0.1%
30.5%
Show abstract

Deep neural networks (DNNs) have been applied to a variety of regulatory genomics tasks. For interpretability, attribution methods are employed to provide importance scores for each nucleotide in a given sequence. However, even with state-of-the-art DNNs, there is no guarantee that these methods can recover interpretable, biological representations. Here we perform systematic experiments on synthetic genomic data to raise awareness of this issue. We find that deeper networks have better generalization performance, but attribution methods recover less interpretable representations. Then, we show training methods promoting robustness - including regularization, injecting random noise into the data, and adversarial training - significantly improve interpretability of DNNs, especially for smaller datasets.

17
Accurate spatial quantification in computational pathology with multiple instance learning

Gao, Z.; Mao, A.; Dong, Y.; Wu, J.; Liu, J.; Wang, C.; He, K.; Gong, T.; Li, C.; Crispin-Ortuzar, M.

2024-04-26 pathology 10.1101/2024.04.25.24306364 medRxiv
Top 0.1%
27.1%
Show abstract

Spatial quantification is a critical step in most computational pathology tasks, from guiding pathologists to areas of clinical interest to discovering tissue phenotypes behind novel biomarkers. To circumvent the need for manual annotations, modern computational pathology methods have favoured multiple-instance learning approaches that can accurately predict whole-slide image labels, albeit at the expense of losing their spatial awareness. We prove mathematically that a model using instance-level aggregation could achieve superior spatial quantification without compromising on whole-slide image prediction performance. We then introduce a superpatch-based measurable multiple instance learning method, SMMILe, and evaluate it across 6 cancer types, 3 highly diverse classification tasks, and 8 datasets involving 3,850 whole-slide images. We benchmark SMMILe against 9 existing methods, and show that in all cases SMMILe matches or exceeds state-of-the-art whole-slide image classification performance while simultaneously achieving outstanding spatial quantification.

18
TaHL-PTM: Post-Translational Modification Prediction in Proteins via Target-Hooked Discriminative Fine-Tuning of Decoder-only Protein Language Models

Prasain, B.; Pratyush, P.; Schulze, S.; KC, D. B.

2026-08-27 bioinformatics 10.64898/2026.08.24.746791 medRxiv
Top 0.1%
27.0%
Show abstract

Post-translational modifications (PTMs) regulate protein function, making accurate residue-level PTM prediction essential for understanding cellular mechanisms and disease pathways. While decoder-only protein language models (PLMs) pretrained with the causal language modeling (CLM) objective have driven breakthroughs across various bioinformatics tasks, their potential for PTM prediction remains largely underexplored. CLM-based PLMs that rely on Byte-Pair Encoding (BPE) for tokenization, such as ProtGPT2, introduce intra-token label collision by merging multiple amino acids with conflicting labels into a single token, creating a major bottleneck for residue-level tasks. To overcome this, we propose TaHL-PTM (Target-Hooked Low-rank adaptation for PTM prediction), a novel framework that integrates target-hooked tokenization with site-directed discriminative LoRA fine-tuning. Target-hooked tokenization constrains tokenization around the candidate residue using dedicated marker tokens to eliminate intra-token label collision while preserving the surrounding sequence context, whereas the proposed discriminative objective repurposes the standard generative CLM objective for residue-level PTM classification by directly optimizing the separation between modified and unmodified sites. We benchmark TaHL-PTM across six distinct PTM tasks on ProtGPT2 and ProGen2 models. TaHL-PTM consistently improves MCC, with the largest gain of up to +0.11 for tyrosine phosphorylation (0.34 to 0.45), alongside improvements in F1, AUROC, and AUPR. Performance gains are more pronounced for collision-affected samples, validating the effectiveness of target-hooked tokenization, while consistent improvements across both BPE-based and per-residue-based causal PLMs demonstrate that the proposed framework generalizes across models with different pretraining tokenization schemes.

19
Generation of synthetic scRNA-seq-like transcriptomes using a generative adversarial network from RNA-seq data

Ruan, D.; Armstrong, S. S.

2025-10-10 bioinformatics 10.1101/2025.10.09.681449 medRxiv
Top 0.1%
26.9%
Show abstract

Next-generation sequencing (NGS) technologies have become integral for high-throughput transcriptomic studies. Among these, single-cell RNA sequencing (scRNA-seq) is especially valuable for quantifying gene expression at the individual cell level, enabling the identification of rare cell populations and cellular differentiation pathways. However, the high cost of scRNA-seq often limits its broader application. Bulk RNA sequencing (RNA-seq) provides a more affordable alternative but lacks the single-cell resolution needed to elucidate cellular heterogeneity. Here, we present a cycle-consistent generative adversarial network (cycleGAN) approach to generate synthetic single-cell-like transcriptomes from bulk RNA-seq data. By adversarially training two sets of generators and discriminators, our framework attempts to learn the relationship between bulk and single-cell transcriptome distributions. Although this approach does not replace real scRNA-seq experiments, it can be a usefult tool to generate synthetic single-cell-like data for preliminary exploratory investigations and other machine learning applications. We further discuss the performance, limitations, and ethical considerations of our method.

20
EvoStructCLIP: A Mutation-Centered Multimodal Embedding Model for CAGI7 Variant Effect Prediction

Chung, K.; Lee, J.; Kim, Y.; Lee, J.; Park, J.; Lee, H.

2026-03-04 bioinformatics 10.64898/2026.03.02.707336 medRxiv
Top 0.1%
26.7%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWWe present EvoStructCLIP, a mutation-centered multimodal embedding model that integrates local 3D structural windows and evolutionary constraints to predict missense variant effects. EvoStructCLIP combines two encoders: a structure voxel encoder derived from AlphaFold residue neighborhoods and an MSA-based evolutionary encoder. It aligns the modalities through CLIP-style contrastive learning, with FuseMix regularization and an auxiliary pathogenicity loss trained on 153,787 ClinVar variants. Evaluations using lightweight regressors demonstrate that EvoStructCLIP embeddings capture highly transferable predictive signals across diverse phenotypes, including gene-specific functional readouts of BRCA1, KCNQ4, and PTEN/TPMT. This transferability is further supported in the CAGI7 blind competition setting, where models generalized to predicting different gene-specific readouts for BARD1, FGFR, and TSC2 without target-specific retraining and achieved competitive performance across heterogeneous biological tasks.