Back

Structure

Elsevier BV

Preprints posted in the last 30 days, ranked by how well they match Structure's content profile, based on 193 papers previously published here. The average preprint has a 0.10% match score for this journal, so anything above that is already an above-average fit.

1
Computational Structural Analysis of POLG Variants R627Q and W748S with Model-Variability Controls

Friedl, A.; Manst, D.

2026-08-27 biophysics 10.64898/2026.08.25.747108 medRxiv
Top 0.1%
15.1%
Show abstract

Background: Comparisons between independently predicted wild-type and missense-variant protein structures can generate mechanistic hypotheses, but small apparent differences may reflect model-selection variability rather than mutation-specific effects. Methods: Human mitochondrial DNA polymerase gamma (POLG; UniProt P54098) variants p.Arg627Gln (R627Q) and p.Trp748Ser (W748S) were evaluated using five AlphaFold2-PTM network-model outputs per condition generated with one random seed under matched ColabFold settings. Ten pairwise wild type comparisons at each site described between-network model-selection variability. Variant effects were summarized across five within-network wild-type-versus-variant comparisons using rotation-invariant local C-alpha pair distances and local displacement after global and local alignment. Because these comparison designs differ, the wild-type distribution was used as context rather than a mutation-effect null. Wild-type cryo-EM structure 9GGF was used for contact and interface mapping. Experimental A467T and G848S structures 9GGE and 9GGC provided contextual benchmarks. Results: R627Q measurements fell within the range of between-network wild-type differences: its median mean local pair-distance change was 0.170 angstrom, compared with a wild-type median of 0.170 angstrom, and its locally aligned displacement was 0.265 versus 0.248 angstrom. W748S showed higher median values (0.168 versus 0.132 angstrom for pair-distance change; 0.236 versus 0.182 angstrom for locally aligned displacement), but the ranges overlapped and the comparison-design asymmetry precluded a calibrated mutation-effect percentile. Experimental A467T and G848S comparisons produced local changes of similar magnitude. In 9GGF, R627 and W748 directly shared a local microenvironment, with a minimum heavy-atom distance of 3.53 angstrom. R627 also formed short polar-contact candidates with D629 and D743, whereas W748 occupied a hydrophobic packing environment containing Y622 and F750. Both sites were more than 18 angstrom from nucleic acid, more than 30 angstrom from POLG2, and more than 33 angstrom from PZL-A in a ligand-bound structure. Conclusions: Available AlphaFold2 comparisons do not establish a mutation-specific structural deformation for either variant. Experimental-structure mapping supports testable physicochemical hypotheses involving a shared R627-W748 microenvironment - loss of an arginine-centered polar network for R627Q and disruption of a buried aromatic environment for W748S - but not direct DNA, POLG2, or PZL-A contact mechanisms. Matched control substitutions and independent seeds are required to calibrate small mutation-associated structural deltas.

2
Cryo-EM Structure of a Triazole alpha-Conotoxin GI Mimetic Bound to the Muscle-Type Nicotinic Acetylcholine Receptor

Shepperson, O.; Capper, M.; Holdship, C.; Melling, O.; Wade, N.; Malone, M.; Arnott, K.; Morgan, D.; Piggot, T.; Morcom, T.; Connah, J.; Windeln, L.; Timperley, C.; Frey, J.; Green, C.; Koehnke, J.; Essex, J.; Jamieson, A.

2026-09-01 biochemistry 10.64898/2026.08.31.748223 medRxiv
Top 0.2%
7.8%
Show abstract

Disulfide-rich peptides possess exceptional potency and selectivity but are often limited by the instability and synthetic challenges associated with native disulfide bonds. Here, we report the design, synthesis, pharmacological evaluation, and structural characterisation of triazole-based peptidomimetics of the -GI conotoxin, a selective antagonist of the muscle-type nicotinic acetylcholine receptor (nAChR). A series of 1,4- and 1,5-disubstituted triazole analogues were prepared entirely on resin using CuAAC and RuAAC chemistry to replace the native Cys3/13 disulfide bridge. Functional evaluation against human muscle nAChRs revealed that 1,5-triazole analogues retained low-nanomolar potency, with the lead mimetic exhibiting activity comparable to native -GI. Cryo-electron microscopy of the lead compound bound to the muscle-type nAChR provided the first structure of a disulfide-isostere peptidomimetic in complex with a membrane receptor. The structure demonstrates that the 1,5-triazole reproduces the native peptide fold with high fidelity while contributing receptor-facing interactions not available to the native disulfide bridge. Molecular dynamics simulations further revealed conserved hydration networks and similar conformational sampling between the native peptide and lead mimetic. Together, these findings establish triazoles as effective disulfide surrogates and provide a structural framework for the rational design of stabilised conotoxin therapeutics.

3
Structural basis of endogenous lipid recognition and G protein selectivity in GPR119

Kim, D.; Ranjbar, M.; Raskovalov, A.; Song, P.; Salve, J.; Katritch, V.; Cherezov, V.

2026-08-28 molecular biology 10.64898/2026.08.27.747701 medRxiv
Top 0.2%
7.8%
Show abstract

G protein-coupled receptors (GPCRs) often engage multiple intracellular transducers, yet the structural basis by which endogenous ligands influence G protein selectivity remains poorly understood. GPR119 is a lipid-activated receptor expressed in pancreatic {beta}-cells and enteroendocrine L-cells, where it regulates glucose-dependent insulin and incretin secretion, making it a promising therapeutic target for metabolic disease. Here, we present cryo-electron microscopy structures of GPR119 bound to its endogenous lipid agonist oleoylethanolamide (OEA) in complex with Gs and Gq proteins. These structures reveal that OEA occupies a deeply buried orthosteric pocket but adopts distinct conformations in the two signaling states. Structural and functional analyses further identify an extended TM5 helix that stabilizes the receptor-Gs interface and acts as a key determinant of G-protein subtype selectivity. Together, these findings provide mechanistic insights into endogenous lipid-driven multi-transducer signaling and establish a structural framework for the development of pathway-selective GPR119 therapeutics.

4
Fine tuning energy metabolism in skeletal muscle: Discovery of a novel autoinhibitory mechanism in the N-terminal extension of AMPKγ3

Ovens, A. J.; Khabib, M. N. H.; Yu, D.; Ling, N. X. Y.; Smiles, W. J.; Hoque, A.; Ann Onda, D.; Poblete Goycoolea, A. C.; Cao, M.; Zhang, G. X. Y.; Turner, B. R.; Doughty, L.; Ang, C.-S.; Horne, C. R.; Scott, J. W.; Sakamoto, K.; Parker, M. W.; Kemp, B. E.; Galic, S.; Oakhill, J. S.; Langendorf, C. G.

2026-08-19 biochemistry 10.64898/2026.08.16.744724 medRxiv
Top 0.2%
6.7%
Show abstract

AMP-activated protein kinase (AMPK) regulates metabolism in response to metabolic stress that includes stimulating glucose uptake in skeletal muscle independently of the canonical insulin signalling pathway, positioning it as an attractive therapeutic target for insulin resistance and type 2 diabetes mellitus (T2DM). AMPK is an {beta}{gamma} heterotrimer, with multiple isoforms for each subunit enabling the formation of 12 different complexes with distinct tissue expression profiles. Among these, the 2{beta}2{gamma}3 complex is predominantly expressed in skeletal muscle, the major site of glucose disposal and a highly desirable therapeutic target for T2DM. Here, we characterise the functional role of a unique, 182 residue N-terminal extension (NTE) within {gamma}3 subunit. Deletion of the {gamma}3-NTE from 2{beta}2{gamma}3 complex increases basal AMPK activity without affecting activation by AMP or pharmacological AMPK activators, demonstrating the {gamma}3-NTE performs an autoinhibitory function. Using complementary biophysical techniques, including hydrogen-deuterium exchange-mass spectrometry, surface plasmon resonance, chemical crosslinking and co-pulldowns, we identified a 39-residue sequence in the {gamma}3-NTE (residues 129-168), that directly interacts with the C-helix of the AMPK kinase domain small lobe, a key regulatory element in many protein kinases. Using AlphaFold3, we probe the interaction predicted to take place between a {gamma}3-NTE -helix ({gamma}3-iHelix; [~]T142-E154) and the C-helix in the 2{beta}2{gamma}3 complex. These findings provide the groundwork for developing novel T2DM therapies that target AMPK activation selectively in skeletal muscle involving reversal of the {gamma}3 autoinhibition.

5
In-Cell Protein Crystallization via a Locally Flexible 24-mer Assembly Precursor

Abe, S.; Tanaka, J.; Kikuchi, K.; Furuta, T.; Aizawa, Y.; Tanaka, Y.; Yokoyama, T.; Kanamaru, S.; Kobayashi, R.; Ueno, T.

2026-08-28 biophysics 10.64898/2026.08.25.746898 medRxiv
Top 0.3%
6.6%
Show abstract

In-cell protein crystallization (ICPC) produces ordered protein crystals within living cells, but the mechanisms used by proteins to acquire long-range crystalline order in the cellular environment remains poorly understood. Here, we define the assembly pathway of CipB, a crystalline inclusion protein from Photorhabdus luminescens. CipB crystals formed in cells dissolve under mild acidic conditions into a predominant 24-mer species, supporting a model in which an in-cell crystal is built from a discrete 24-mer assembly precursor rather than through direct packing of smaller oligomeric states. Structural analysis of recrystallized CipB shows that the same 24-mer architecture packs into a body-centered cubic lattice, consistent with the lattice observed for the in-cell crystals. Cryo-EM and molecular dynamics analyses indicate that the 24-mer assembly precursor preserves its overall architecture while retaining local conformational flexibility at the N-terminal and surface-loop regions. Mutation analyses further link the N-terminal region to the formation of the 24-mer precursor and surface residues to lattice assembly. These observations support a stepwise crystallization model in which N-terminal flexibility facilitates the formation of an assembly-competent 24-mer precursor, whereas defined hydrophobic surface contacts subsequently organize these precursors into a long-range-ordered lattice.

6
Inferring protein ensembles directly from NOESY spectra

Coles, M.

2026-08-23 biophysics 10.64898/2026.08.20.745893 medRxiv
Top 0.3%
6.2%
Show abstract

Solution NMR spectroscopy provides atomistic measurements of proteins in a native-like biophysical state. Because these measurements are ensemble averages, it also has the potential to report on conformational diversity. However, conventional NMR structure determination typically converts experimental observables into restraints for molecular dynamics, which encode information on the mean structure but do not retain information on the underlying conformational distribution. Ensemble selection has long been proposed as an alternative, whereby experimental observables are compared directly with candidate conformers generated independently of the measurements. This allows population distributions to be inferred from the data. However, few such methods have incorporated NOESY - the richest source of structural information in protein NMR - data, due to challenges in the quantitative comparison of experimental and back-calculated spectra. To address this challenge, we previously introduced the CoMAND method, demonstrating that quantitative agreement is practical for NOESY spectra with bespoke heteronuclear editing schemes. Here we extend this approach into a framework for direct inference of protein ensembles within a flexible ensemble-selection architecture incorporating multiple classes of NMR observables. We introduce a quantitative scoring framework for comparing experimental and back-calculated observables and combine it with regularized ensemble selection and Monte Carlo simulated annealing. Integration with the OpenMM molecular dynamics engine allows conformational pools to be generated using established molecular simulation methods. Applied to human ubiquitin, the resulting ensemble provides simultaneous agreement with NOESY, residual dipolar coupling and scalar coupling data while retaining conformational diversity supported by experiment.

7
Molecular architecture of colossal surface layers from hyperthermophilic archaea

Caspy, I.; Cvirkaite-Krupovic, V.; van Dorst, S.; von Kuegelgen, A.; Ford, Z.; Alva, V.; Krupovic, M.; Bharat, T. A. M.

2026-08-21 microbiology 10.64898/2026.08.21.746120 medRxiv
Top 0.3%
6.1%
Show abstract

Surface layers (S-layers) are paracrystalline protein lattices that form the outermost layer of the cell envelope in most archaea, providing structural support, protecting against external insults, and co-ordinating interactions with their environment. Despite their widespread occurrence, the molecular and structural details of S-layer architecture in hyperthermophilic archaea remain largely unknown. Here, we report the structure and cellular architecture of the S-layer from the hyperthermophilic archaeon Pyrobaculum arsenaticum by combining in situ electron cryotomography with single-particle electron cryomicroscopy, AlphaFold modelling, and peptide-fingerprinting mass spectrometry. We show that the S-layer is formed by an uncharacterised 292-kDa S-layer protein (SLP) extending 37 nm from the cytoplasmic membrane, making it, to our knowledge, the largest SLP structurally characterised to date. This SLP has a remarkable multidomain architecture comprising 19 immunoglobulin-like domains, 14 canonical and five non-canonical, organised into a lattice-forming core, a stalk, and a unique crown domain that stabilise the S-layer. Comparative genomic analyses unearthed homologous colossal SLP candidates across Thermoproteota, indicating that this distinctive architecture is conserved across diverse archaeal lineages. Together, our findings provide a structural framework for understanding the cell-surface organisation in hyperthermophilic archaea and suggest that these colossal S-layers represent a specialised adaptation to life at high temperatures.

8
CD36 phosphorylation alters the thrombospondin binding site and reduces internal cavity accessibility and volume

Ghojoghi, G.; Chemtob, S.; Lubell, W. D.; Ong, H.; Meneksedag Erol, D.

2026-09-01 biophysics 10.64898/2026.08.25.747030 medRxiv
Top 0.3%
5.6%
Show abstract

The cluster of differentiation 36 (CD36) is a membrane protein with broad physiological roles in health and disease, and its function is regulated in part by phosphorylation. Experimental evidence shows that phosphorylation of Thr92 reduces CD36 affinity for thrombospondin-1 (TSP-1), binding of which initiates antiangiogenic signaling, whereas phosphorylation of Ser237 decreases CD36-mediated fatty acid uptake, with implications for energy metabolism. However, the only available crystal structure of CD36 lacks phosphorylation, and the molecular mechanisms by which phosphorylation regulates CD36 function remain largely unknown. This study provides an atomically detailed computational characterization of CD36 in unphosphorylated and dual phosphorylated states, using molecular dynamics simulations with a total sampling time of 30 microseconds in combination with Markov state models. We present, to our knowledge, the first evidence of a cryptic pocket on CD36 surface that is formed by phosphorylation. This cryptic surface pocket and a loop spanning residues 121-131 form a high affinity binding site for TSP-1 derived ligands, shifting their binding away from the canonical site. We propose that this altered binding provides a molecular basis for the disruption of antiangiogenic signaling upon CD36 phosphorylation. Additionally, our data indicate that, phosphorylation increases helicity and compaction within the helix-loop region spanning residues 296-331, narrowing one of the entrances to the internal cavity and reducing its overall volume. These conformational changes provide a potential mechanistic explanation for the decrease in fatty acid uptake upon CD36 phosphorylation. Our findings provide structural insights that may inform the future design of CD36 modulators and emphasize the importance of targeting phosphorylation induced CD36 conformations in angiogenic and metabolic diseases.

9
Heparan sulfate selectively inhibits the collagenase activity of matrix metalloproteinase 13

Hao, H.; Su, G.; Liu, J.; Xu, D.

2026-08-24 biochemistry 10.64898/2026.08.21.746339 medRxiv
Top 0.3%
5.5%
Show abstract

Matrix metalloproteinase 13 (MMP13) is a zinc-dependent protease that plays key roles in extracellular matrix remodeling. Like several other MMPs, MMP13 has been shown to interact with heparan sulfate (HS), a highly sulfated glycosaminoglycan found at the cell surface and in the extracellular matrix, but the significance of the interaction remains unknown. Here we report that while zymogen and mature forms of MMP13 both bind HS with high affinity, their interactions with HS display markedly different characteristics in terms of preferred HS structure and binding kinetics. By structure-guided mutagenesis, we identified a large HS-binding site of MMP13 consists of 10 residues in the hemopexin domain, 3 residues in the catalytic domain, and 2 residues in the linker region. While these basic residues participate in binding to both zymogen and mature forms of MMP13, the relative contribution of many residues differs substantially between the two forms, which likely contributes to their distinct HS-binding characteristics. Binding of HS to mature MMP13 resulted in selective inhibition of the collagenase activity of MMP13 in a length- and sulfation-dependent manner, but the binding had no effect on degradation of non-collagen substrates. Mechanistically, the inhibitory effect of HS likely results from reduced interdomain flexibility after binding of HS, and/or HS-induced dimerization of MMP13. In sum, our study establishes HS as a multifaceted regulator of MMP13 activity, and discovers that the HS-binding site of MMP13 is a novel exosite that can be targeted to inhibits its collagenase activity.

10
A structural census links penultimate-residue class to N-terminal burial in human protein assemblies

Chang, Y.-H.

2026-09-01 biochemistry 10.64898/2026.08.31.748389 medRxiv
Top 0.3%
5.5%
Show abstract

Initiator-methionine excision is among the earliest protein modifications, yet its relationship to assembly geometry is unknown. Burial of the mature first residue was measured across 7,246 deposited human biological assemblies (22,291 chain-level observations; 1,191 proteins). Among 1,143 analyzable proteins, termini in MetAP-permissive penultimate-residue sequence classes were less often interface-engaged than termini in MetAP-nonpermissive classes (37.4% versus 47.4%; adjusted odds ratio 0.65, p = 7.2e-4). Curated processing annotations did not show a corresponding burial difference, and correlated residue properties preclude attributing the sequence-class association specifically to iMet removal. The analysis identified 264 interface-engaged MetAP-permissive candidates concentrated in cellular machines. In a fully recomputed conformer scan of deeply buried proteasome positions, modeled methionine accommodation was less favorable than at observed-methionine controls (median overlap -0.30 versus -1.12 angstrom, p = 0.0049), although most scoreable sites permitted a nonoverlapping placement. The census therefore reveals a graded structural constraint - not universal steric failure - and prioritizes complexes in which altered packing, assembly kinetics, lipidation or N-terminal methylation can be tested.

11
Visualizing Reaction Pathways via Reciprocal Space Kinetic Decomposition

Grunewald, L.; Meszaros, P.; Westenhoff, S.

2026-08-20 biophysics 10.64898/2026.08.17.745189 medRxiv
Top 0.4%
5.4%
Show abstract

Time-resolved serial crystallography (TR-SX) has emerged as a powerful method for capturing ultrafast structural dynamics in proteins. TR-SX continues to produce remarkable studies, revealing previously unobserved transient states and providing deeper insights into processes such as drug targeting, DNA repair, and photosynthesis. However, extracting weak structural signals from noisy time-resolved datasets remains a major challenge. Robust computational methods are therefore required to isolate the signals associated with the underlying transient states. Importantly, this should be performed in reciprocal space to preserve compatibility with established downstream structure refinement workflows. Here, we introduce a framework for kinetic decomposition directly in reciprocal space that enables separation of kinetically distinct structural states. The method decomposes crystallographic data according to a predefined kinetic model, improving the recovery of weak transient signals and enhancing mechanistic interpretation from limited time-resolved datasets. We validate the framework using simulated data based on a previously published time-resolved crystallography study and demonstrate its application to a new TR-SX dataset comprising 17 time points. We show that the method separates the reciprocal space signatures of four intermediates by incorporating kinetic information from a predefined reaction model. This establishes a workflow for extracting kinetic states directly from time-resolved X-ray diffraction data that can be seamlessly integrated into existing crystallographic structure-determination pipelines.

12
WaterFlow: Prediction of Ordered Water Molecule Positions on Protein Structures

Srivastava, V.; Mai, H.; Collins, M.; HOLTON, J. M.; Wall, M.; Wankowicz, S. A.

2026-08-27 biophysics 10.64898/2026.08.26.747373 medRxiv
Top 0.4%
5.2%
Show abstract

Ordered water molecules mediate many protein functions, including stability, ligand binding, and catalysis. Predicting their positions with sub-angstrom accuracy would support protein design, binding affinity prediction, and automated model building in X-ray crystallography and cryo-EM. However, water molecule prediction lags behind protein and other molecule structure predictions. Here, we introduce WaterFlow, a flow-matching-based generator model and confidence model for predicting the positions of ordered water molecules in protein structures. WaterFlow outperforms the existing state of the art at every precision level. We demonstrate that WaterFlow can accurately predict ground truth modeled water molecules, including those around protein-ligand interactions and on predicted structures. We also show that WaterFlow predictions fit well directly to experimental data, and therefore propose that it may be used for both prediction and modeling water molecules. This includes novel predictions that are often associated with positive electron difference density, meaning the model places water molecules at sites the original structure depositions omitted. We use this improved model to address the data constraint. By mapping the Pareto front of achievable accuracy of water molecule prediction, alongside analysis of different training data schemas, we quantified the trade-off between data quantity and data quality, demonstrating that the diversity of high-quality structures is limiting the possible results. Overall, WaterFlow predicts ordered water to serve as a solvent module for structure-based drug design and for water molecule placement during crystallographic refinement.

13
Structural basis of K+/H+ antiport in YcgO and its inhibition by unphosphorylated PtsN

Srivastava, A.; Athreya, A.; Patidar, Y.; Singh, V.; Sardesai, A. A.; Penmatsa, A.

2026-08-18 biochemistry 10.64898/2026.08.14.744764 medRxiv
Top 0.4%
5.1%
Show abstract

Cation-proton antiporters (CPAs) are vital for the maintenance of ionic homeostasis and normal physiology among diverse cell types. Despite recent insights into K+/H+ exchange transporters, the diversity in their structural organization and regulatory mechanisms of K+-specific CPAs are minimally understood. Here, we explore the architecture of an E. coli CPA1 K+/H+ antiporter, YcgO and its inhibition by the unphosphorylated form of PtsN, the terminal protein of a regulatory phosphorelay, using cryoEM structures at 3.4 [A] and 3.2 [A] resolution, respectively. Homodimeric YcgO bound to K+ ions in the occluded conformation, harbors additional linked cytosolic domains, RCK and CorC, to regulate the movement of the transport helices within the YcgO dimer. These domains are the sites of interaction and efflux inhibition by unphosphorylated PtsN, which interacts with the CorC domains with high affinity and allosterically augments inhibitory interactions of CorC with transport helices of YcgO. Inhibition is relieved leading to constitutive activation, upon disrupting the CorC-transport conduit interface. This study illuminates the structural basis of K+ efflux mediated through regulation of a K+/H+ antiporter in E. coli and related prokaryotes via a metabolic network involving a regulatory phosphorelay.

14
Makeshift: a lightweight software for accessing and analyzing NMR data and protein dynamics

El Nesr, G.; Wayment-Steele, H. K.

2026-08-20 biophysics 10.64898/2026.08.17.745346 medRxiv
Top 0.4%
5.1%
Show abstract

Nuclear magnetic resonance (NMR) spectroscopy yields rich residue-level information on biomolecular dynamics and chemical environments, two frontiers for quantitative predictive methods in biochemistry. Decades of data are publicly archived in the Biological Magnetic Resonance Data Bank (BMRB)1, yet in practice, this information remains difficult to access and interpret at scale and within computational workflows. Here we present makeshift, an open-source Python package for accessing, curating, and analyzing NMR datasets. Users can readily retrieve and parse BMRB entries and perform essential analyses such as chemical shift re-referencing, secondary structure propensity prediction, and interpretation of relaxation datasets for dynamics. We re-implemented several widely-used NMR data calculations which were not open-source or available in Python and validated our implementations against the original implementations. By integrating data access, processing, and analysis into a single Python interface, makeshift lowers the barrier for reproducible, scalable analysis and machine learning applications using biomolecular NMR data.

15
A multi-scale structural and biophysical atlas of TCR-peptide-HLA recognition dynamics

Zhang, S.; Long, Y.; Wang, T.; Zhong, Q.; Li, J.; Fu, L.

2026-08-26 molecular biology 10.64898/2026.08.19.745737 medRxiv
Top 0.5%
4.8%
Show abstract

Dynamic interactions between T cell receptor (TCR) and peptide-human leukocyte antigen (pHLA) complexes are central to peptide-specific immune recognition, influencing T cell activation and immune responses. While structural biology has provided valuable static structures of TCR-pHLA complexes, systematic datasets capturing their dynamic and interaction patterns remain limited. Here, we present DynaTPH, a curated structural dynamics dataset of human TCR-pHLA complexes. DynaTPH integrates TCR-pHLA structures, covering both HLA class I and class II complexes, and extends these static structural resources with standardized molecular dynamics simulations and derived biophysical properties. Through a multi-stage filtering procedure, we identified 256 representative complexes and performed standardized all-atom molecular dynamics simulations for each system, corresponding to a cumulative simulation time of 38.4 s. The dataset includes static structures, trajectories, corresponding frames, and derived physicochemical properties, including hydrogen bonds, intermolecular contacts, solvent accessibility, and backbone flexibility. By capturing the conformational flexibility and dynamic interaction patterns across diverse TCR-pHLA interfaces, DynaTPH extends static structural resources with multidimensional biophysical information. This dataset enables systematic investigation of TCR-pHLA recognition dynamics and supports applications in TCR engineering, vaccine design, and immune tolerance research and artificial intelligence-driven computational immunology.

16
Pi-Ensemble: Sequence-guided generation of interpolated protein conformational ensembles

Nadeem, H.; Kleiman, D. E.; Zhou, Y.; Leakey, A. D. B.; Shukla, D.

2026-08-18 biophysics 10.64898/2026.08.12.744498 medRxiv
Top 0.5%
4.4%
Show abstract

Proteins are critical biomolecular machines that populate ensembles of interconverting conformations. Many biological processes depend on transitions between metastable states. Although molecular dynamics (MD) simulations provide a physically grounded route to characterize these motions, routine sampling of large-scale conformational transitions remains computationally demanding. Recent advances in protein structure prediction have created new opportunities for ensemble generation, but many existing approaches require noising inputs, task-specific training, supervised fitting on extensive MD data, or experimentally-informed restraints. Here, we introduce Pi-Ensemble (Predicting Interpolated Ensemble), a sequence-guided framework for generating protein conformational ensembles interpolating between two structural anchor states. Unlike previous methods, Pi-Ensemble alternately leverages inverse-folding and structure-prediction models to propose intermediate conformations between known protein states, generating diverse ensembles without additional training. We evaluate Pi-Ensemble across diverse protein systems, including enzymes, transporters, receptors, and benchmark cases with reference MD simulations or experimental Double Electron-Electron Resonance (DEER) data. Pi-Ensemble recovers physically plausible intermediate conformations, captures transition pathways observed in large-scale MD simulations, and generates structures consistent with experimental distance distributions. Furthermore, Pi-Ensemble-generated conformations provide effective starting seeds for parallel MD simulations, improving conformational exploration and accelerating convergence relative to simulations initiated only from endpoint structures. These results establish sequence-guided structural interpolation as a practical strategy for probing protein conformational landscapes. By generating diverse and physically reasonable conformational proposals without long-timescale MD or model retraining, Pi-Ensemble provides an extensible framework for studying protein flexibility, guiding adaptive sampling, and accelerating mechanistic investigations of protein function.

17
CryoForge: A Self-Correcting Agent for Cryo-EM Model Building That Learns When to Act and When to Stop

Feng, W.; Jiang, y.; Sun, F.; Yang, J.; Gao, X.; Zhang, F.; Han, R.

2026-08-20 bioinformatics 10.64898/2026.08.15.745007 medRxiv
Top 0.5%
4.4%
Show abstract

Automated atomic model building has accelerated cryo-EM structure determination, but different builders leave distinct residual error profiles requiring expert inspection. The post-building challenge is to decide which local interpretations are sufficiently supported by experimental evidence to be retained, corrected or rejected. Here we introduce CryoForge, an evidence-gated post-builder agent that separates repair proposal from repair acceptance. Rule and learning-based components identify candidate regions and prioritize legal actions, whereas an independent evidence gate evaluates each edit using map and half-map support, stereochemistry, connectivity and local structural context. Supported edits are retained; unsupported or conflicting modifications are rejected, rolled back, stopped or escalated for expert review. Across a resolution-stratified benchmark, 84.9% of 26,153 released trajectories yielded standard validated improvements and 3.7% yielded low-confidence partial improvements, with no quality-degrading edit retained in the final promoted models. Relative to rule-only control, learned prioritization reduced non-improving candidates and harmful actions while preserving global structural stability. External evaluations using an alternative initializer, same-team automated/manual-assisted challenge submissions and three recently released complex assemblies showed that CryoForge adapts to distinct residual error phenotypes and performs bounded, evidence-supported correction without uncontrolled remodeling. CryoForge provides a builder-independent, scalable and auditable correction layer between automated model generation and expert structural interpretation.

18
Characterizing the interaction of a type VII-secreted antimycobacterial toxin with its small helical partner proteins

Lee, E.; Bowran, K.; Boardman, E.; Palmer, T.

2026-08-25 microbiology 10.64898/2026.08.24.746431 medRxiv
Top 0.5%
4.3%
Show abstract

The type VII secretion system (T7SS) is a membrane-embedded protein export pathway found in mycobacteria and Gram-positive bacteria. Recently it was shown that Mycobacterium abscessus uses its ESX-4 variant of the T7SS to secrete a toxin, EatA, which targets arabinogalactan present in the mycobacterial cell envelope. Prior to its export, EatA forms a complex with a pair of small proteins from the WXG100 family, TapA1 and TapA2. Here we investigated a structural model of the EatA N-terminal domain in complex with TapA1 and TapA2 using site-directed mutagenesis and bacterial 2-hybrid assays. Our results are consistent with the three proteins forming a stacked bundle of alpha-helices. Structural modelling also predicted an interaction of the EatA-TapA1-TapA2 complex with EsxT-EsxU, a second pair of WXG100-family proteins that are likely required for the mechanistic operation of ESX-4. Whilst we could demonstrate a potential interaction between TapA2 and EsxT by bacterial 2-hybrid analysis, we were not able to purify a complex of all five proteins.

19
Structure of a dodecameric double-ferritin-fold protein from an Asgard archaeon

Remeeva, A.; Anuchina, A.; Dashevskii, D.; Kurkin, T.; Semenov, O.; Mishin, A.; Osipov, S.; Li, G.; Shishkin, P.; Shuvaev, Y.; Mikhailov, A.; Kuznetsova, E.; Natarov, I.; Nikolaev, A.; Sudarev, V.; Vlasov, A.; Borshchevskiy, V.; Rogachev, A.; Gushchin, I.

2026-08-26 biophysics 10.64898/2026.08.25.747088 medRxiv
Top 0.6%
4.1%
Show abstract

Ferritins are ubiquitous iron homeostasis proteins found across the tree of life that form conserved 24-subunit cages with octahedral (4-3-2) symmetry. New types of ferritins and ferritin-like proteins are being continuously discovered, such as mini-bacterioferritins, which form smaller shells of 12 subunits, and double-ferritin-fold proteins, which act as ferroxidases but do not form shells. Here, we describe double-ferritin-fold proteins from Asgard archaea, dubbed dFTNs, and determine Cryo-EM structure of a representative from Candidatus Heimdallarchaeum endolithica. The protein forms a dodecameric shell with tetrahedral (2-3) symmetry. N-terminal (NTD) and C-terminal (CTD) domains are bridged by an ordered linker and are related by two-fold rotational pseudosymmetry. C-terminal -helix (helix E) that forms the four-fold channel in classic ferritins is repositioned to be the helix 2 out of 5 ferritin domain -helices in dFTN, with two such helices from NTD and two helices from CTD forming a pseudo-four-fold symmetry structural element. Four three-fold channels are formed by NTDs, and four other such channels are formed by CTDs. The overall arrangement of dFTN ferritin domains is similar to that of protomers in classic ferritin shells. Altogether, our findings expand the range of known ferritin family proteins and provide insight into Asgard archaea iron metabolism.

20
ARCHER: Amortized cross-specimen pose estimation for cryo-electron microscopy

Nguyen, N.; Pham, B.

2026-08-21 biophysics 10.64898/2026.08.21.746234 medRxiv
Top 0.6%
4.0%
Show abstract

Single-particle cryo-electron microscopy (cryo-EM) pose estimation is traditionally solved anew for each dataset, where iterative refinement is done from scratch while the estimator learns to store the molecule in its weights. In this work, we show that pose inference is a generalizable, specimen-agnostic operation when conditioned explicitly on a reference volume. We introduce ARCHER, an amortized contrastive classifier that models the pose posterior over a discrete rotation grid. Trained across a variety of protein structures, it operates zero-shot without retraining per structure. This transferability is grounded in Fourier-space information mechanics, where all specimen dependence is captured by the reference structure's power spectrum and spatial extent. ARCHER achieves a median angular error of 5.0{degrees} on 100 held-out test structures and 2.5{degrees} on experimental particles, matching dedicated estimators within 0.16[A] in 3D reconstruction. Crucially, downstream conformational signal is preserved. The leading conformational coordinate correlates at 0.97 with deposited benchmarks, faithfully reconstructing free-energy basins and mobile domains. These results overall demonstrate that cryo-EM pose estimation can be generalized across different structures.