Back

mAbs

Informa UK Limited

Preprints posted in the last 7 days, ranked by how well they match mAbs's content profile, based on 32 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
Site Specific Fluorescent Labeling via SpyTag SpyCatcher for Rapid Hybridoma Screening in Semi-Solid Medium

Guo, A.; Wei, M.; Wu, J.; Li, X.; Jiang, B.

2026-08-31 immunology 10.64898/2026.08.21.746134 medRxiv
Top 0.1%
9.1%
Show abstract

Hybridoma screening in semi-solid medium typically employs antigens labeled with visible fluorophores (e.g., FITC, AF488) to enable single-step identification of antibody-secreting clones. However, conventional chemical conjugation via NHS-esters or isothiocyanate groups frequently modifies lysine residues located within epitopes, potentially abrogating antibody recognition of these critical regions. Here, we describe a SpyTag SpyCatcher-based site-specific labeling strategy that circumvents epitope damage during semi-solid medium screening. A 16-amino-acid SpyTag was genetically fused to the C-terminus of the target antigen, enabling covalent conjugation to an sfGFP SpyCatcher fluorescent probe. In semi-solid medium supplemented with SpyTag-antigen and sfGFPSpyCatcher, positive hybridoma clones were readily identified by distinct fluorescent halos, whereas negative clones showed no detectable signal. Notably, the site-specific method yielded a significantly higher frequency of fluorescence-positive clones compared to the conventional AF488-labeled antigen method, suggesting that epitope preservation enhances screening recovery. Furthermore, this approach did not impair hybridoma growth or final clone positivity, offering a simple, rapid, and epitope-compatible method for monoclonal antibody screening.

2
The first OpenBind release: An open experimental structure-affinity dataset and benchmark for structure-based AI

Nelen, J.; Khan, O.; Adams, E.; Aschenbrenner, J. C.; Thompson, W.; Ebrahim, A.; Capkin, E.; Vallee, C.; OpenBind, ; Shotton, E. J.; Griffen, E. J.; Chodera, J. D.; Deane, C. M.; von Delft, F.; AlQuraishi, M.; Imrie, F.

2026-09-01 bioinformatics 10.64898/2026.08.27.747600 medRxiv
Top 0.2%
2.4%
Show abstract

High-quality experimental datasets that link protein-ligand structures with binding affinity data are essential for developing and evaluating structure-based machine learning methods. To help address this need, we established OpenBind as an open-science initiative to generate large-scale experimental datasets for structure-based AI and molecular discovery. Here, we describe the first public OpenBind release, which, to the best of our knowledge, is the largest public single-target experimental structure-affinity dataset. The dataset focuses on enteroviral 2A protease, comprising 925 crystallographic binding events from 699 compounds and associated affinity measurements for 601 compounds. It combines structures from an initial fragment screen and follow-on molecules, together with affinity data, linking experimentally determined protein-ligand binding modes to biophysical measurements within a coherent antiviral discovery campaign. We used this dataset to evaluate protein-ligand structure prediction, binding-affinity prediction, and virtual screening using representative structure-based methods, including docking and cofolding. This exposed several challenges that are central to practical structure-based modelling: docking performance depends strongly on binding-pocket conformation, poses are difficult to rank, and structure-based affinity prediction remains challenging. Fine-tuning OpenFold3-p2 on the fragment-screen structures substantially improved pose prediction and virtual screening for related follow-on compounds, demonstrating how early-stage experimental structures can support target-specific model adaptation.

3
FlexiTAC enables controllable PROTAC linker generation across diverse structural settings using a Bayesian flow network with posterior guidance

Li, Y.; Zhao, Y.; Zhou, L.; Huang, C.; Xu, Q.; Chen, Y.; Qin, Z.; Fan, K.; Yang, J.; Cao, D.

2026-08-30 bioinformatics 10.64898/2026.08.26.747172 medRxiv
Top 0.3%
1.7%
Show abstract

Linker chemistry and conformation are central determinants of PROTAC activity, shaping ternary-complex geometry, cooperativity, target-lysine presentation and cellular permeability. Existing linker generators often lack explicit control over linker flexibility, require predefined attachment sites and linker lengths, or produce structures that demand substantial geometric correction, limiting their utility in practical PROTAC design. Here we introduce FlexiTAC, a Bayesian flow network that jointly generates linker atom types and coordinates from the warhead and E3-ligase-ligand contexts. We also assemble PROTAC-3D, a quality-controlled collection of 63,554 component-resolved PROTAC structures for model training, and PROTAC-Bench, which covers molecular quality, fragment preservation, geometric fidelity, conformational stability, fragment awareness, rediscovery and sampling efficiency. Compared to the best 3D baseline models, FlexiTAC improves validity by 12.0-12.7% and achieves the highest PoseBusters pass rate of 79.5%-80.0%. A differentiable guidance module shifted generated linkers along a conformational ensemble-derived rigidity axis without retraining the generator. In silico case studies further show that the model can accept crystal-derived, redocked or predicted structural inputs. Together, FlexiTAC, PROTAC-3D and PROTAC-Bench establish an integrated and reproducible framework for data-driven PROTAC linker design, combining controllable structure-conditioned generation with standardized training data and evaluation protocols. This framework expands the linker chemical and conformational space accessible to computational exploration, provides a foundation for future method development and enables the systematic generation of structure-conditioned linker designs with tunable conformational flexibility.

4
Design and characterization of broadly protective influenza A(H3N2) vaccine candidates using protein language models

Howard, V. R.; Allen, J. D.; Thomas, M. H.; Sautto, G. A.; Ross, T. M.; Georgiev, I. S.

2026-08-31 immunology 10.64898/2026.08.26.747087 medRxiv
Top 0.3%
1.4%
Show abstract

Seasonal influenza A viruses cause significant global morbidity each year. Although vaccination remains the primary preventive strategy, effectiveness is often reduced by antigenic drift. This challenge is particularly pronounced for influenza A(H3N2), which has required eight vaccine updates over the past decade. Here, we present a computational framework to engineer broadly reactive influenza A(H3N2) vaccines, using protein language models to generate novel hemagglutinin (HA) sequences and a machine learning model to predict antigenic distance from circulating strains. In a proof-of-concept study, seven HA candidates designed using sequence data from 2013-2018 were evaluated in mice against contemporary and subsequently circulating viruses. Two candidates elicited protective levels of reactive antibodies, robust H3-specific antibody-secreting cell responses, and cross-neutralization against contemporary clades and drifted 2019-2020 strains. These findings demonstrate that an integrated generation-selection strategy can enhance vaccine coverage across current and future A(H3N2) seasons and may be applicable to other influenza subtypes.

5
Chemi-Proteome Language Attention Network Empowers Fragment-Based Ligand Interactome and Binding Sites Discovery with Evidence

Liao, B.; He, J.; zhao, M.; Cui, X.; Cui, Y.; Dong, C.; Sun, H.; Zhang, L.; Zhang, J.

2026-08-30 bioinformatics 10.64898/2026.08.26.747036 medRxiv
Top 0.5%
0.6%
Show abstract

Deep learning has accelerated drug discovery, yet most existing models are trained using in vitro affinity datasets and consequently remain disconnected from the cellular context in which functional ligand-protein interactions occur. This limitation hinders the ability to reflect the complexity of native interactomes and characterize biological responses to molecular perturbation. Here we introduce C-PLANK (Chemi-Proteome Language Attention NetworK), a deep learning framework trained on fragment-protein interactions profiled directly in living cells using fully functionalized fragment (FFF) chemoproteomics. C-PLANK combines physicochemical embeddings with a bilinear attention network (BAN) to model both global cellular context and local residue-atom interactions, generating interpretable interaction fingerprints. Particularly, C-PLANK incorporates Cellular Interaction State Index (CISI), a systems-level evidential metric that contextualizes the biological plausibility of each predicted interaction against the global cellular interaction landscape. Across 431 ligand interactomes curated from eight independent chemoproteomic studies, C-PLANK consistently outperformed current state-of-the-art interaction prediction frameworks under both random and cold-protein evaluation settings. The inferred interaction fingerprints aligned with orthogonal evidence from structure-based pocket predictions, co-crystal structures, and cellular binding-site annotations. C-PLANK further generalized to unseen ligands. In a cellular target-focused discovery campaign, C-PLANK identified a previously unrecognized ligand that was subsequently advanced into an active chemical probe acting as a SIRT3 agonist in cellular assays. By learning directly from cellular chemoproteomics, C-PLANK moves beyond isolated interaction prediction toward cellular interaction-state modelling, establishing a computational foundation for future digital-twin frameworks in drug discovery.

6
From Prompt to Provenance: BloClaw, a Capability-Gated AI4S Workstation for Auditable Computational Biology

qin, y.; Pang, J.; Zhang, X.

2026-09-01 bioinformatics 10.64898/2026.08.26.747436 medRxiv
Top 0.7%
0.5%
Show abstract

Scientific agents can produce plausible answers while remaining unable to establish whether the computation behind an answer is executable, recoverable, or reproducible. We present BloClaw, an AI4S workstation built around a simple principle: a scientific agent should know what it can do, show how it did it, and state what remains unvalidated. Each capability declares an execution state, input constraints, dependencies, expected outputs, and scientific limitations. Natural-language requests are translated into structured tasks, validated against this registry, executed through scientific tools, and recorded in a provenance-aware Living Lab Notebook. The system is designed to detect invalid inputs, failed tool calls, missing dependencies, and remote timeouts, and to route them to repair, retry, or escalation. The implemented and tested scope comprises RDKit-based molecular property and rule screening, protein structure analysis, docking-pose inspection, 3D visualization, and structured reporting. We demonstrate the workflow on a PubChem-retrieved osimertinib structure and a supplied 6LU7 docking artifact: the former yields deterministic descriptors (molecular weight 499.619 Da, cLogP 4.5098, TPSA 87.55 A^2), while the latter contains 2,387 protein ATOM records, 309 residues, and nine pose records. These examples are workflow demonstrations, not efficacy or affinity studies. Beyond retrospective prediction, the manuscript specifies a prior-minimized constructive mode in which a desired function is compiled into explicit physical, chemical, and systems constraints, candidate mechanisms are simulated, and observations are reintroduced for calibration and falsification; this is a proposed extension rather than a result of the present case studies. We describe an evaluation protocol that compares BloClaw with a standard single-agent workflow and fixed-script execution using task completion, scientific correctness, recovery success, provenance completeness, reproducibility, human review time, latency, and cost. This manuscript reports the system design, verified capability boundary, deterministic software artifacts, and a reproducible evaluation protocol; it does not claim benchmark improvements before those experiments are run. BloClaw is an execution and accountability layer for AI-assisted research, complementing expert review and experimental validation rather than replacing them.

7
Cryo-EM Structure of a Triazole alpha-Conotoxin GI Mimetic Bound to the Muscle-Type Nicotinic Acetylcholine Receptor

Shepperson, O.; Capper, M.; Holdship, C.; Melling, O.; Wade, N.; Malone, M.; Arnott, K.; Morgan, D.; Piggot, T.; Morcom, T.; Connah, J.; Windeln, L.; Timperley, C.; Frey, J.; Green, C.; Koehnke, J.; Essex, J.; Jamieson, A.

2026-09-01 biochemistry 10.64898/2026.08.31.748223 medRxiv
Top 0.8%
0.3%
Show abstract

Disulfide-rich peptides possess exceptional potency and selectivity but are often limited by the instability and synthetic challenges associated with native disulfide bonds. Here, we report the design, synthesis, pharmacological evaluation, and structural characterisation of triazole-based peptidomimetics of the -GI conotoxin, a selective antagonist of the muscle-type nicotinic acetylcholine receptor (nAChR). A series of 1,4- and 1,5-disubstituted triazole analogues were prepared entirely on resin using CuAAC and RuAAC chemistry to replace the native Cys3/13 disulfide bridge. Functional evaluation against human muscle nAChRs revealed that 1,5-triazole analogues retained low-nanomolar potency, with the lead mimetic exhibiting activity comparable to native -GI. Cryo-electron microscopy of the lead compound bound to the muscle-type nAChR provided the first structure of a disulfide-isostere peptidomimetic in complex with a membrane receptor. The structure demonstrates that the 1,5-triazole reproduces the native peptide fold with high fidelity while contributing receptor-facing interactions not available to the native disulfide bridge. Molecular dynamics simulations further revealed conserved hydration networks and similar conformational sampling between the native peptide and lead mimetic. Together, these findings establish triazoles as effective disulfide surrogates and provide a structural framework for the rational design of stabilised conotoxin therapeutics.

8
Cryo-EM Structure of Duck Secretory IgM Reveals a Conserved Pentameric Assembly with Avian-Specific Features at Molecular Interfaces

Schneider, R. M.; Liu, Q.; Stadtmueller, B. M.

2026-08-30 immunology 10.64898/2026.08.26.747385 medRxiv
Top 1%
0.1%
Show abstract

IgM is the most ancient antibody isotype, playing an important role in both circulatory and mucosal immune responses across vertebrates, yet structural characterization of its polymeric forms is limited outside of mammals. Here, we report the cryo-electron microscopy structure of mallard duck secretory (S) IgM at 3.37-[A] resolution. The structure revealed a pentameric core globally similar to human SIgM, supporting the view that pentameric IgM is subject to strong evolutionary constraints. However, compared to mammalian structures, we observed species-specific differences at molecular interfaces. Surface plasmon resonance binding assays characterizing secretory component (SC)-IgM interactions supported structural observations and, when compared to IgA binding, revealed isotype-specific contributions from the avian SC N-terminal extension. Together, these findings establish a comparative structural framework for polymeric IgM across vertebrates and provide insight into how avian SIgM-specific features may support mucosal immunity in birds.

9
Isolation and characterisation of Nipah virus neutralising candidate therapeutic monoclonal antibodies from an mRNA-immunised pig

Pedrera, M.; Pipatpadungsin, N.; Kobasa, D.; Elrefaey, A. M. E.; Holzer, B.; McLean, R. K.; Warner, B.; Vendramelli, R.; Thakur, N.; Stass, R.; Hayes, J. W. P.; Medfai, L.; Sealy, J. E.; Crossley, S.; Schwartz, J. C.; Munir, D.; Mwangi, W.; Bailey, D.; Truong, T.; Tchilian, E.; Pickering, B.; Bowden, T. A.; Graham, S. P.

2026-08-30 immunology 10.64898/2026.08.28.745669 medRxiv
Top 1%
0.1%
Show abstract

Nipah virus (NiV) is a highly pathogenic zoonotic paramyxovirus with epidemic potential. Despite the threat NiV poses, no therapeutics are licensed to treat infection. Studies have shown that monoclonal antibodies (mAb) can protect animals against NiV and the related Hendra virus (HeV). The best studied mAb, m102.4, has been used to treat infected patients on a compassionate basis, and has entered clinical trials. However, there is a need to define additional mAbs with therapeutic potential, which could be combined with m102.4 to improve neutralising potency and breadth. Here, we isolated five high affinity mAbs from an mRNA immunised pig, which bound the G glycoprotein derived from NiV Malaysia strain (NiV-M), and one of which (mAb A2) also bound HeV G. Aligned with this, all mAbs neutralised NiV-M pseudovirus but only mAb A2 neutralised pseudovirus representing the NiV Bangladesh (NiV-B) strain. mAb A2 and the most potent NiV-M neutralising mAb, C1, showed minimal competition with each other and m102.4, suggesting recognition of non-overlapping epitopes. Single-particle cryogenic electron microscopy of the NiV-M G receptor binding domain complexed to A1 and C2 Fab fragments revealed distinct epitopes that did not overlap with the receptor-binding site, targeted by m102.4, suggesting action through steric impedance of receptor binding or interference downstream of receptor engagement. Inoculation of mAb A2 to hamsters did not provide complete protection against NiV-B challenge (60% survival), however, a split dose of mAb A2 and m102.4 provided the same protection as m102.4 alone (100% survival). Collectively, these data demonstrate the potential of the porcine model for isolation of therapeutic candidate mAbs, which contribute both to our understanding of the NiV G antigenic landscape, and the development of mAb combinations, that exert complementary mechanisms of neutralisation, for therapeutic intervention.

10
In silico engineered multitarget-directed ligands for the polypharmaceutical treatment of PTEN loss of function endometrial adenocarcinoma

Delara, R.; Mujumdar, V.; Zhang, Q.; Dryden, H.; Crane, E.; Brown, J.; Naumann, W.; Puechl, A.; Foureau, D.; Sha, W.; LeGrand, J.; Yang, H.-T.; Dykema, K.; Yada, B.; McHale, C. C.; Maddeboina, K.; Pal, D.; Durden, D. L.

2026-08-31 cancer biology 10.64898/2026.08.28.747865 medRxiv
Top 1%
0.1%
Show abstract

To combat refractory diseases, such as cancer, multitarget-directed ligands (MTDLs) have become an emerging area of research to exploit synthetic lethality (SL) relationships associated with drug resistance. Herein, we present the in silico design of MTDLs for the polypharmaceutical treatment of endometrial adenocarcinoma (EAC) and our discovery of a novel SL in EAC; PTEN loss of function (LOF) and the inhibition of CDK9. We used high-resolution x-ray crystallographic data to chemically engineer, LCI133, to inhibit CDK9, CDK4/6-and AURKA/B kinases. PTEN LOF in EAC results in augmented deregulated transcription and a massive increase in nascent RNA, a phenotype which encodes a high level of apoptotic sensitivity to LCI133 and CDK9 inhibitors. Treatment with LCI133 results in a rapid decline nose-dive in global nRNA, MYC nRNA levels and TS elongation (TE) in PTEN LOF EAC. PTEN LOF is necessary and sufficient to confer sensitivity of EAC cells to LCI133 and other CDK9 inhibitors.

11
ChemIntelligence Enables Antibody-Free, Ultra-Low-Input Profiling of Lysine Lactylation and Diverse Acyl-Proteomes

Shao, C.; He, Z.; Yuan, Q.; Giurcoiu, V.-G.; He, X.; Cao, X.; Huang, H.; Zhang, Y.; Zhang, Y.; Wang, D.; Jiang, Q.; Guo, Z.; Hao, H.; Wilhelm, M.; Ye, H.

2026-08-31 biochemistry 10.64898/2026.08.28.746934 medRxiv
Top 2%
0.1%
Show abstract

Lysine acylations, including lactylation (Klac), are pivotal regulators of cellular physiology. However, their analysis is currently bottlenecked by antibody enrichment strategies that suffer from sequence bias and require milligram-scale protein inputs, severely precluding the profiling of scarce clinical biopsies and rare cell populations. Here we present ChemIntelligence, an acyl-NHS chemistry-empowered derivatization strategy that rapidly generates unprecedented acylation-specific spectral libraries, exemplified by over 2.5x10^9 human Klac peptides, enabling cross-species reference atlases. Integrated with Prosit-based rescoring, these libraries substantially increase Klac identifications across diverse proteomic datasets. Leveraging this spectral resource, we devised ChemIntelligence Scope, a reproducible, multiplexed parallel reaction monitoring (PRM) platform that quantifies hundreds of Klac peptides per injection from as little as ~200 ng of cell lysates, clinical biopsies, and even true single cells - revealing functional Klac signatures inaccessible to conventional methods. The ChemIntelligence pipeline also extends seamlessly to lysine nicotinylation, underscoring its broad adaptability for discovering and profiling new acylations. Together, these chemical and computational advances establish a scalable, antibody-free framework for acyl-proteome mapping that overcomes input constraints and enables deep functional insights from otherwise intractable biological samples.

12
RegimeFormer: A Large Protein Model of Global Perturbation Regimes

Ma, S.; Chai, Y.; Wu, Y.; Zhang, Q.; Yuan, Y.; Zhao, K.; Chen, Z.; Wang, H.; Cao, S.; Yu, X.; Han, X.; Liu, Y.; Liu, Y.; Zhu, T.; Tao, D.

2026-08-30 bioinformatics 10.64898/2026.08.26.747182 medRxiv
Top 2%
0.1%
Show abstract

Protein language models organize sequence and structure at scale, but a global representation of how proteins respond to mutation remains lacking. We present RegimeFormer, a large protein perturbation model coupled to RegimeAtlas, constructed by harmonizing and indexing 202,556,313 non-redundant protein sequences across the tree of life. A diversity-preserving one-million-protein subset provides the high-resolution training and inference layer, with 995,995 proteins yielding residue-level summaries across 407,048,356 residues and substitution-specific predictions available on demand. Across experimental deep mutational scanning, molecular benchmarks, structural confidence and evolutionary constraint, RegimeFormer identifies reproducible protein-level perturbation regimes that organize residue fragility, adaptability and predictive uncertainty. Regime conditioning improves substitution-specific prediction, with the largest relative gains under unseen-protein, unseen-family and low-homology evaluation. RegimeFormer-derived molecular priors further improve downstream transcriptomic and drug-response modelling. Together, RegimeFormer and RegimeAtlas provide a scalable framework for mapping, predicting and querying protein perturbation landscapes across global sequence space.

13
Development and Optimization of 111In-Dinutuximab-IRDye800, a Dual-Modality Intraoperative Molecular Imaging Agent for Pediatric Neuroblastoma Resection

Yip, C. Y.; Rosenblum, L. T.; Pant, A.; Kahler-Quesada, A.; Chagantipati, B.; Sever, R.; Grano-Mickelsen, B.; Li, B.; Cortez, A. G.; Latoche, J. D.; Day, K. E.; Rigatti, L.; Nedrow, J. R.; Edwards, B. W.; Kohanbash, G.; Malek, M. M.

2026-08-31 cancer biology 10.64898/2026.08.28.747876 medRxiv
Top 2%
0.1%
Show abstract

Rationale: Neuroblastoma is a devastating pediatric malignancy, for which surgical resection is a key factor in long-term survival. However, there are significant challenges in its resection, particularly in high-risk disease, as neuroblastoma encases surrounding critical structures, is often difficult to distinguish from desmoplastic or scar tissue, and can carry occult deposits of disease not readily identified on preoperative imaging or intraoperative visualization. Building on the principles of fluorescent and radio-guided surgery, in combination with the known overexpression of GD2 in neuroblastoma, we sought to develop and optimize 111In-Dinutuximab-IRDye800, a dual-modality GD2-targeted intraoperative molecular imaging agent, for use in pediatric neuroblastoma to help enhance patient safety while facilitating a more complete resection. Methods: Dinutuximab was conjugated to IRDye800 and DTPA, then radiolabeled with Indium-111 to yield 111In-Dinutuximab-IRDye800. Optimization occurred through ELISA assay to assess binding affinity, fluorescence intensity analysis to determine the optimal fluorescent degree of labeling, and phototoxicity testing through flow cytometry. Rodent models of neuroblastoma were then generated through injection of SK-N-BE(2) human neuroblastoma cells into the left adrenal glands of nude mice or RNU rats. A series of fluorescent and gamma biodistributions was performed, varying the dose, timing, and specific activity of the tracer. Tumor and organ uptake of the tracer was compared with one- or two-way ANOVA as appropriate, with Sidaks multiple comparison test to compare tumor uptake to individual organs. Once optimization was complete, a clinically significant events study modeled after human clinical trials was performed to evaluate the in vivo capabilities of 111In-Dinutuximab-IRDye800. Results: Increased ratios of IRDye800 per antibody led to decreased binding affinity for GD2 and was associated with formulation instability without significant return on fluorescence intensity. Specific activity of the tracer was not found to impact overall biodistribution of the tracer. A 45-50 microgram dose of 111In-Dinutuximab-IRDye800 with ratios around 1 DTPA and 1-1.5 IRDye800 per antibody imaged 4 days after tracer administration was found to be the optimal combination that maximized detectable tumor-specific signal. In the clinically significant events study mirroring human IMI clinical trials, fluorescent guidance identified additional malignant lesions not originally detected under white light in 64% of rodents. Conclusions: 111In-Dinutuximab-IRDye800 is a dual-modality GD2-targeted intraoperative imaging agent that is well-poised for clinical translation. As it preserves tumor specificity, yields clinically meaningful radiofluorescent signal, and is well-tolerated without adverse events after optimization was completed, it carries the potential to positively impact the safety and completeness of neuroblastoma resection.

14
Every Cure Knowledge Graph: A Unified Biomedical Knowledge Graph for Drug Repurposing

Kaniewski, P.; Carter, E. K.; Rhodes, D.; Lim, E. M.; Li, J.; Vergine, J.; Matentzoglu, N.; Schaper, K.; Reilly, J.; Sundar, S.; Vijnck, L.; Sharp, E.; Alfonso, N.; Ford, A.; Stepanenko, A.; Hempstead, C.; Brokmeier, P.; Bizon, C.; Tropsha, A.; Haendel, M. A.; Fajgenbaum, D. C.; Lancashire, L.

2026-08-31 bioinformatics 10.64898/2026.08.26.747253 medRxiv
Top 2%
0.1%
Show abstract

Identifying causal connections between existing drugs and mechanistic profiles of diseases is a foundational step for effective drug repurposing. Although knowledge graphs (KGs) are highly suited for consolidating biomedical databases and tracking these connections, a single biomedical KG is constrained by its ingestion pipeline and knowledge sources. While different biomedical KGs could be complementary if combined, efforts to combine them into a unified and more comprehensive KG are hindered by lack of interoperability and poor provenance. To address those issues, we present EC-KG, a Biolink Model-compatible KG for computational drug repurposing. EC-KG is an interoperable, provenance-first KG which integrates RTX-KG2, ROBOKOP, and PrimeKG at the network-level, encapsulating over 7 million nodes and 81 million edges from 95 primary data sources. EC-KG has improved coverage of core biomedical entities such as drugs, targets, and diseases relevant to drug repurposing vs source graphs, and captures complex biomedical mechanisms within its topology. We demonstrate that the network unification in EC-KG leads to emergence of novel, mechanistically relevant pathways which are disconnected in the underlying constituent networks and show its applications in method development, benchmarking and predictive drug repurposing applications. EC-KG has already been successfully used in drug repurposing research to surface Botulinum Toxin A as a candidate to treat Major Depressive Disorder, as well as to validate repurposing of Lenalidomide and Dexamethasone for a subgroup of patients with Rosai-Dorfman Disease.

15
Aerolysin enables modular, non-genetic functionalization of living cell surfaces

Lemmex, A. C.; Pawlak, M. R.; Gordon, W. R.

2026-08-31 biochemistry 10.64898/2026.08.28.746739 medRxiv
Top 2%
0.1%
Show abstract

Methods for installing synthetic functions on living cell surfaces provide powerful approaches for imaging, sensing, and manipulating cell behavior, but many require genetic modification of the target cell or chemical modification of the plasma membrane. Here, we repurpose the glycosylphosphatidylinositol-anchored protein (GPI-AP)-binding toxin aerolysin as a modular chassis for non-genetic cell-surface functionalization. We show that a non-cytotoxic, monomeric aerolysin mutant retains high-affinity and GPI-AP-dependent cell binding when genetically fused to diverse protein cargos. Fluorescent protein-aerolysin fusions robustly label multiple cell types and remain predominantly associated with the cell surface for at least 24 h, in contrast to wheat germ agglutinin, which is extensively internalized. Aerolysin can also be equipped with SpyTag/SpyCatcher to enable modular assembly with independently expressed protein cargos. Importantly, aerolysin supports functional rather than solely optical modification of the cell surface: fusion to the proximity-labeling enzyme APEX2 enables extracellular protein biotinylation, while fusion to HUH endonuclease tags enables covalent attachment of synthetic DNA to living cells. Using this latter architecture, we developed a DNA hairpin sensor that converts cell-surface nuclease activity into a fluorescent signal and distinguishes cells with different levels of extracellular nuclease activity. Together, these results establish non-cytotoxic aerolysin as a genetically encoded, soluble adapter for installing proteins, enzymes, and programmable nucleic acids onto living cells without modification of the target-cell genome.

16
Accurate and efficient prediction of protein conformations with ProtMonomer

Si, Y.; Zhang, S.; Chen, L.

2026-08-31 molecular biology 10.64898/2026.08.28.747824 medRxiv
Top 2%
0.1%
Show abstract

Deep learning-based protein structure prediction methods that leverage evolutionary information from multiple sequence alignments (MSAs), exemplified by AlphaFold2, have achieved remarkable accuracy. However, existing methods still struggle to predict challenging proteins, particularly those with novel folds or limited evolutionary information, and to recover alternative conformational states. Here we show that structure prediction models trained under different MSA-depth distributions corresponding to different levels of evolutionary information exhibit complementary generalization behaviors, and that a model trained on a mixture of these distributions can combine their complementary generalization strengths. Building on this insight, we developed ProtMonomer, a deep learning framework trained on MSA-depth distributions representing a broad range of evolutionary information levels to improve structure prediction. Across benchmarks comprising CASP15 targets, non-redundant experimentally determined structures, orphan proteins, and short peptides, ProtMonomer performed comparably to or better than leading methods, including AlphaFold2 and AlphaFold3, with particularly strong performance on challenging targets. For fold-switching proteins, ProtMonomer also recovered alternative conformational states more accurately than AlphaFold2 and AlphaFold3 across diverse homologous sequence sampling strategies. In addition to improving predictive accuracy, ProtMonomer substantially reduced inference cost through an efficient architecture, enabling high-throughput applications. Together, these findings provide insights into the generalization of evolution-informed structure prediction models and support ProtMonomer as an accurate and efficient framework for protein structure prediction.

17
PathFold: Predicting the Entire Protein Folding Pathway from Protein Sequence Alone

Zhang, Z.; Ibtehaz, N.; Kagaya, Y.; Xu, Z.; Punuru, P.; Kihara, D.

2026-09-01 bioinformatics 10.64898/2026.08.26.747321 medRxiv
Top 2%
0.1%
Show abstract

Recent advances in protein structure prediction, exemplified by AlphaFold, have largely addressed the determination of static structures, one aspect of the protein folding problem. However, predicting folding pathways, by which proteins reach their native states, remains a significant challenge. Here, we present PathFold, a deep learning framework that predicts protein folding pathways directly from sequence information. PathFold leverages an AlphaFold-based module to extract structural information from the sequence and generates a progressive folding trajectory from an extended conformation using a diffusion model. By modeling the full trajectory, it enables prediction of folding intermediates and transition pathways, analogous to those observed in steered molecular dynamics (SMD) simulations. The predicted pathways reveal well-defined intermediates and sequential folding events, and show agreement with experimental folding data, including measured {Phi}-values.

18
A Metabolic Labeling Strategy for Tracking Protein Synthesis in Complex Biological Systems

Bu, Y. J.; Nyandwi, S. P.; De Lima Alves, F.; Tennakoon, R.; Stamm, T. V.; Schneider, D. J.; Eddenden, A.; Ma, T. W. Y.; Chun, Y.-j.; Peng, H.; Miller, J. M.; Wheeler, A. R.; Yuzwa, S.; Nitz, M.; Cui, H.

2026-09-01 molecular biology 10.64898/2026.08.30.747940 medRxiv
Top 3%
0.1%
Show abstract

Protein synthesis supports most biological processes. In the brain in particular, protein synthesis plays a critical role in physiological and pathological states. Here, we describe Tellurophene-Alkyne Cycloaddition-mediated Amino acid Tagging (TeACAT), a versatile strategy for fast, facile, and flexible tagging of newly synthesized proteins in mice. TeACAT is based on metabolic incorporation of the non-canonical amino acid TePhe into proteins by the endogenous protein synthesis machinery. Due to their high similarity, TePhe can efficiently replace canonical Phe without dietary or genetic manipulation. The subsequent bio-orthogonal reaction of TePhe with either fluorescent dyes or affinity handles enables both visualization and affinity enrichment of proteins synthesized during TePhe exposure. TeACAT is compatible with immunofluorescence for cell-type specific visualization of protein synthesis with subcellular resolution and can be used in conjunction with routine proteomics to identify and quantify newly synthesized proteins. Robust incorporation into the mouse proteome was observed on the scale of hours to days, allowing the interrogation of various biological processes. In summary, TeACAT enables the visualization and quantification of protein synthesis with minimal perturbation for biological discoveries.

19
ECG-based longitudinal risk prediction across diseases and organ systems

ye, y.; Zeng, Z.; Tian, X.; Yuan, Z.; Wang, J.; Zhu, Y.

2026-09-02 health informatics 10.64898/2026.08.29.26361697 medRxiv
Top 3%
0.0%
Show abstract

Artificial intelligence applied to routine electrocardiograms (ECGs) has largely focused on detecting existing disease or predicting individual cardiovascular outcomes. Whether ECGs can support prediction of multiple future diseases across organ systems remains unclear. We developed ECG-RISK, a multitask survival model for 67 incident three-character ICD-10 endpoints using ECG waveforms, demographic characteristics and routinely collected laboratory data from 86,673 MIMIC-IV patients. Discrimination was highest for heart, brain, kidney and lung endpoints, with organ-level C-indices ranging from 0.796 to 0.825, whereas liver and pancreatic endpoints showed lower discrimination. The ECG-only model achieved strong discrimination across most endpoints, whereas the incremental improvement gained by incorporating ECG and laboratory inputs beyond demographic information varied substantially across endpoints. Across the nine exploratory aggregated outcomes, Kaplan Meier curves showed clear separation among model-score tertiles. Discrimination was highest for dementia (C-index, 0.891) and heart failure (C-index, 0.857). These findings support the feasibility of ECG-based longitudinal risk prediction across multiple diseases. External validation and competing-risk analyses are required to assess generalisability and clinical utility.

20
Burden of fatigue in compensated chronic liver disease: findings from the multinational a:GAP Study

Choudhuri, G.; Akhundova-Unadkat, G.; Naidoo, N.; Morales-Castillo, M.; Guillaume, X.; Duijnhoven, R. G.; Safaei, A.; Swain, M. G.

2026-09-02 gastroenterology 10.64898/2026.08.28.26361618 medRxiv
Top 3%
0.0%
Show abstract

Background & Aims: Fatigue is a central symptom of chronic liver disease (CLD), substantially impacting health-related quality of life (HRQoL). This study aimed to further understand CLD symptomatology, including fatigue, and its impact on HRQoL from a patient perspective. Methods: Abbott Global Assessment of Patients unmet needs (aGAP) was a multinational, cross-sectional survey in adults with compensated CLD in China, India and Mexico, conducted between July and November 2024. Adult participants who self-reported that they had physician-diagnosed CLD and were experiencing fatigue completed a quantitative survey to assess symptom burden and included three HRQoL patient-reported outcome (PRO) questionnaires (Patient-Reported Outcomes Measurement Information System [PROMIS]-29+2, Work Productivity and Activity Impairment - Specific Health Problem version 2.0 [WPAI: SHP], Multidimensional Fatigue Inventory [MFI]). Results: Overall, 505 participants (China: 200; Mexico: 105; India: 200) completed the study. Participants reported that their CLD-related fatigue sometimes, often or always affected their self-esteem/confidence (45.1%) and ability to maintain or acquire new employment (38.6%). Most participants reported moderate (51.3%) or serious (26.9%) fatigue, with 33.5% experiencing fatigue every day or almost every day. Many participants felt their social life was negatively impacted by their fatigue (47.3%) and that there were related financial difficulties (53.9%). Use of validated PRO tools demonstrated severe fatigue (MFI: overall mean [SD] 13.9 [3.4] general fatigue and 13.4 [3.6] physical fatigue) as well as substantial levels of work and activity impairment (WPAI: SHP overall mean [SD] 53.0 [26.4]) and high levels of anxiety, pain interference, depression and sleep interference (PROMIS T-scores [≥]54). Conclusions: Fatigue has a substantial impact on HRQoL among adults with CLD across several countries, highlighting a global unmet need for targeted interventions to effectively identify and manage the condition.