Entropy
○ MDPI AG
Preprints posted in the last 90 days, ranked by how well they match Entropy's content profile, based on 21 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Seifer, S.
Show abstract
Progress in quantum computation offers new opportunities for addressing longstanding combinatorial challenges. One such challenge is the Eternity II edge-matching puzzle, consisting of 256 tiles, which has resisted solution despite extensive community effort. The computational complexity of this NP-complete problem exceeds the capacity of current quantum annealing processors but lies within reach of hybrid quantum-classical solvers. Testing a quadratic unconstrained binary optimization (QUBO) model of a puzzle on a D-Wave hybrid solver demonstrates a complete solution only for puzzle instances up to 64 tiles. Simulated quantum annealing fails on this benchmark, whereas an original classical heuristic, "nucleation with deduction", succeeds. To approach the full Eternity II puzzle, I developed a MATLAB package that integrates multiple quantum and classical approaches, including neural-network transformers and gradient-based refinement. A multistage computation pipeline is demonstrated successfully on a puzzle comparable in complexity to Eternity II and with a known solution, based on multiple hybrid optimization steps with both "hard" and "soft" constraint formulations, identification of persistent substructures, and a final classical refinement stage. The resulting optimization problem involves [~]100,000 logical variables and requires partial initialization. Intriguingly, solving this puzzle mirrors the "end game" of protein folding, a process that nature completes in mere fractions of a second, seemingly defying expectations set by the Levinthal paradox. The prospect of predicting protein structure by quantum annealing is reviewed in light of these results.
Sarti, E.; Cazals, F.
Show abstract
The Nobel prize winning program AlphaFold2 computes plausible structures of (well) folded proteins. The main quality assessment is based on the predicted Local Distance Difference Test (pLDDT), a per amino acid confidence score. To enhance quality assessment, we provide novel quantitative measures to identify coherent amino acid (a.a.) stretches along the sequence in terms of pLDDT values. These constructions, grounded in standard techniques from topological data analysis and combinatorics, provide a canonical framework for identifying regions along the protein backbone and analyzing their properties, such as their propensity for disorder and their consistency with a null model. The outcome of our analysis can readily be used to select reliable regions/domains within proteins whose pLDDT values span the entire pLDDT range.
Kringelbach, M. L.; Deco, G.
Show abstract
Brain dynamics can be described in three different convenient mathematical languages, namely connectome harmonics, turbulence and complex harmonics (CHARM). Here we demonstrate that these theoretical frameworks can be rigorously unified, under the functional calculus, as one self-adjoint operator and its single spectral measure. The connectome Laplacian carries that measure; the harmonics are its spectral projections, the turbulence smoothing kernel is its resolvent, and the CHARM form is its unitary propagator. The bridge that makes this exact is a textbook fact: The exponential distance rule, which is the empirical kernel of the turbulence model, is the Greens function of a screened Laplacian, so the local order parameter is the phase field passed through the resolvent of the same operator whose eigenfunctions are the harmonics. A single shared control parameter, the spectral gap, simultaneously yields the cortical hierarchy, the turbulent information cascade and the structured interference the CHARM form measures. This unification makes a strong predictive claim. If the harmonic projections, the turbulence resolvent and the CHARM propagator really are three functions of one operator, then any structural perturbation that re-tunes the operator must move all three signatures in unison and must do so with a single coupling. We test this prediction with a pharmacological perturbation by lysergic acid diethylamide (LSD), which is known to change the emotional state, by empirically perturbing the operator with a 5-HT2A receptor density map and asking whether one scalar coupling can simultaneously predict the multi-scale turbulence shift observed, through the resolvent, and the macroscale harmonic energy redistribution, through first-order Rayleigh-Schrodinger perturbation theory. We found that the two independent functional domains respond in unison to one structural perturbation of one operator. The identity is exact as operator calculus and its purchase on the brain depends on a single load-bearing seam, the degree heterogeneity of the connectome, which we make explicit. We propose that this single-operator structure is the necessary mathematical scaffolding of our Entangled Loop theory.
Song, H.; Hu, G.; Wu, X.; Zhang, X.; Li, J.
Show abstract
Biomolecular condensates are widespread cellular self-assembled structures with essential functions. There are suggestions of condensates formed by different proteins being near criticality. However, systematic investigation of the criticality of condensates is absent, and critical exponents defining their universality class have not been found. Here, using long-time simulations, we show that condensates exhibit typical critical phenomena, including scale-free spatiotemporal correlations, critical slowing down, divergence of correlation length and dynamic scaling. From these scaling behaviors, a set of critical exponents is determined. Based on dynamic critical exponent, diverse condensates can be divided into two distinct universality classes, arising from differences in their molecular components and interaction types.
Truong, Q. H. X.; Truong, X. K.
Show abstract
The emergence of amino acids (AAs) and nucleobases (NBs) across meteorites, interstellar ices, and laboratory shock experiments presents a paradox: why do these specific molecular motifs--a minuscule subset of organic chemistrys combinatorial space--appear repeatedly across diverse environments, in the absence of biological selection? We identify a physical mechanism, prebiotic selection, which biases driven chemical systems toward configurations with high stationary probability p*(x) under sustained entropy flux. The bias is quantified by an information quasi-potential {Phi}I (x) = - ln p*(x), entering the overdamped Langevin dynamics O_FD O_INLINEFIG[Formula 1]C_INLINEFIGM_FD(1)C_FD where {Sigma} is the local entropy production rate (Schnakenberg 1976). {Phi}I is defined self-consistently via the full non-equilibrium stationary density, avoiding the circularity of identifying it with a scalar potential. Two central theorems underlie the framework. Theorem 1 establishes that {nabla}{Sigma} and {nabla}{Phi}I are generically linearly independent off equilibrium, so the dynamics is genuinely two-field. Theorem 2 (structural constraints on single-field gradient dynamics) shows that single-field models on compact manifolds (i) produce yield curves that are at most unimodal under linear driving, and (ii) combine disjoint perturbations additively, giving superlinearity factor S = 1 + O(||{delta} V ||2). The observed superlinear synergy of Ferris et al. (1996) lies far outside this perturbative bound and therefore requires the two-field structure of EOM-IFF; the non-monotonic peak of Blank et al. (2001) is consistent with two-field dynamics and also with single-field dynamics in the unimodal-with-peak case of Theorem 2 part 1, so it does not by itself discriminate. From these results, we: (i) define a formal substrate-minimal criterion for prebiotic selection; (ii) show consistency with the non-monotonic shock-synthesis yield of Blank et al. (2001) (R2 = 0.885, peak at P* = 28.4 {+/-} 1.4 GPa); (iii) show consistency with the superlinear clay-catalysed RNA polymerisation of Ferris et al. (1996) (synergy factor S {approx} 5.75, robust under {+/-}1-nucleotide measurement uncertainty); and (iv) state two further falsifiable predictions awaiting dedicated experimental tests. Every lemma and theorem is accompanied by explicit assumptions, regime of validity, and regime of failure; the frameworks scope is what it claims, not more. Prebiotic selection is identified as a physical process distinct from and prior to biological selection, offering a unified account of chemical convergence in carbon-nitrogen chemistry under sustained entropy flux.
Rossi, A.; Smecca, A.
Show abstract
Two-dimensional accounts of consciousness that distinguish global integration from functional diversity are empirically supported [1,2] but lack a formal phase-structure: they do not specify the nature of the transition between the two regimes, the order parameter that governs it, or the quantitative predictions that follow. We provide this structure. We propose that the dimensionality of the neural correlation structure, operationalised as the Participation Ratio of the covariance eigenspectrum, constitutes a second, independent order parameter D that governs a phase transition distinct from global integration {Phi}. Formalised through a Landau-Ginzburg free energy functional F[{Phi}, D], this transition defines a Redundant Integrated State {Delta} (high {Phi}, low D) in which globally integrated mental function is present but phenomenal experience is absent, a thermodynamic phase, not a point on a continuum. The framework generates three falsifiable predictions absent from prior work: (i) a power-law scaling D* [~] |{Phi} - {Phi}_c|^{nu} with measurable critical exponent{nu} ; (ii) a diverging susceptibility {chi}_D = {partial}D/{partial}{Phi} at the consciousness threshold, quantifiable from perturbational EEG; (iii) an explicit dissociation between MCS and VS patients in the ({Phi}, D) space, with MCS predicted to occupy state {Delta}. These predictions are directly testable with existing methodology and are not generated by any current theory of consciousness.
Ali, A. F.; Inan, N.; Laukkonen, R.; Mikheenko, P.
Show abstract
We develop a theoretical proposal linking vacuum stability and brain dynamics through superconductivity-inspired coherence, symmetry reduction, and the thermodynamic stabilization of low-entropy regimes. We take an unbroken SU(3) structure as a candidate stable residue of the low-temperature vacuum. At the neural level, we formulate a coarse-grained analog in which a two-fluid model with dissipative and coherence-supporting components describes brain dynamics. Specifically, the coherence-supporting component is proposed as a possible basis for the efficient binding and integration required to sustain a stable, unified conscious state. The proposal offers a common geometric language for relating physics and neuroscience with falsifiable signatures in coherence and state-dependent transitions. The main technical contribution is a computational algebraic model of conscious-state dynamics, where neural data are mapped to reconstructed state trajectories. Effective generators are inferred from those trajectories, and the two-fluid split is tested as a Cartan-root decomposition of su(3), with a rank-two commuting sector for coherence-preserving balance and six root directions for state transitions. This structure can be tested on neural data and contrasted with alternative dynamical models.
Park, J.; Smith, C.; Tseng, S. Y.; Guidera, J.; Semenov, A. V.; Smirnov, S.; Frank, L. M.; Pao, G. M.
Show abstract
Complex systems such as brains and other interacting biological and physical processes are difficult to represent because they evolve across many variables, scales, and nonlinear interactions. To capture these multivariate, multiscale interactions we have developed Generative Manifold Networks (GMNs) a machine learning framework consisting of a network of linked dynamical systems. The network is discovered by an interaction function which can focus on causality, shared information, nonlinearity or other metric. Network nodes are low-dimensional data-driven state-space manifolds with generator functions accommodating multiscale dynamics. In contrast to many machine learning approaches GMNs have no latent or randomly initialized variables providing transparent explainability. GMNs generate short term dynamics of chaos on par with echo state networks while outperforming them in long term generation of chaos and neural dynamics, but with a markedly reduced number of dimensions and without sensitive dependence on reservoir parameters or random states. As a result of their holistic, multiscale representation GMNs can learn the complete dynamics of a complex system. We further show that GMNs are universal approximators. GMNs are demonstrated on chaotic dynamics, neural and behavioral recordings of the fruit fly and domestic rat with comparisons to echo state networks and crossformer - a time series transformer. SignificanceA major challenge in machine learning is to model complex systems accurately without losing interpretability. Many methods that succeed in prediction rely on latent variables obscuring mechanistic insight and complicating experimental testing. Generative manifold networks (GMN) construct a network of low-dimensional functional manifolds directly from observed variables with no latent or randomly initialized variables: the model remains transparent and experimentally testable. We prove that GMN are universal approximators showing that high representational power can be achieved without sacrificing explainability. GMN therefore provides a general framework for prediction and simulation in neuroscience and complex systems where unraveling the links between variables in an experimentally testable manner is as important as forecasting their behavior.
Yokoyama, H.; Takeuchi, R.; Shimizu, S.
Show abstract
The primary objective of system neuroscience is to understand the functional mapping and its causation in the dynamics of the brain network. Some experimental and methodological studies suggest that functional modularity and its hierarchical information processing in the brain network are crucial to understanding the functional role of task-specific or state-specific information flow in the brain. However, because most of the established techniques for detecting effective network structures in the neuroscience research field are strongly based on the "Granger causality" perspective, existing causal discovery methods specified for brain network analysis cannot identify the causal hierarchy in the modular network in the brain due to spurious correlation issues and indistinguishability of causal direction under the Gaussianity of observational noise in a linear system. To address the issues, we developed a causal discovery method for synchronous neural dynamics, called the Jacobian-informed linear non-Gaussian acyclic model, "j-VAR-LiNGAM", by incorporating the information of the Jacobian matrix determined from a phase-coupled oscillator model estimated from observed neural data into the VAR-LiNGAM algorithms. The method was validated by showing that it could extract causal ordering in both synthetic data and empirical neural observed data. Moreover, by analyzing the observed neural oscillatory signals obtained from mice and humans, we confirmed that our method identified causally hierarchical structures in the brain, which aligned with the neurophysiological interpretations. These findings suggested that our proposed method can reveal the neural basis of hierarchical information processing in the brain network.
Foster, P. P.; Chhikara, R. S.; Boriek, A. M.
Show abstract
Despite extensive study of cellular mechanisms underlying long-term potentiation, no single specific protein or gene has been identified which encodes an individual unit of information, or memory bit. Indeed, the brain engram remains a knowledge gap. The theory of exclusion led us to cancel one-by-one several unrealistic biological options, suggesting that the explanation resides somewhere else. Superposition of up to concentric 300 myelin layers, spiraled, and highly compacted wrapping a single axon and each wrap could host hundreds to thousands of niches, as memory cells, collectively consisting of a massive array of cells. The disjointed 3D spatial superposition allows storage of charges, nodes not facing from a layer to next. The thickness of a single myelin layer ranges from 7.0 to 20 nm. The dimension scale is approximately the exact dimensions of the charge trap, the tunnel and dielectric also equipping current AI microchips. Stored charges are positive ions, with similar effect whether charges are negative or positive charges creating an electromagnetic field. To write data, following an action potential, this voltage applies to the control gates of the myelin layers producing an ionic charge injection. This causes charges to gain energy and tunnel through the myelin layer across Ranvier nodes, via quantum tunneling, and deep into the concentric myelin multilayers. This is creating an insulated trapping of K+ ions isolated from the system. In a long white matter tract bundle, the near-perfect isolation of millions of axons within compressed myelin wrap-ion channel K+/Na+ systems provides quantum coherence and precision of asynchronous firing property. The injected ionic charges (K+) become physically stuck in traps within the myelin layers. The K+ ions may not move freely, completely trapped after AP ceases. Mirroring a single-bit, single-level-cell, a trapped ionic charge (ions K+) may represent a 1, while an empty cell (absence of K+) represents a 0. The trial-and-error process, with a Bayesian inference which may have also been the core evolution of the learning human brain. Based on selected mathematical equations, we analyzed the general scheme on how deep learning may be embedded in the brain
Konstorum, A.; Xing, J.; Aeron, S.; Kilmer, M.; Kleinstein, S.
Show abstract
Systems-level immune profiling data arising from longitudinal studies of vaccination or infection has an inherent multi-index array structure. While tensor decomposition of such datasets has gained popularity, choosing a rank and trial for a decomposition is not straightforward. We show that taking into account the experimental data model can inspire the development of new metrics to assess the quality of a Non-negative CANDECOMP/PARAFAC (NCPD) decomposition, and can thus be used to choose a rank and trial for the decomposition. Moreover, we show how framing the results via a dictionary learning framework can better enable interpretation of the components of the decomposition.
Xie, J.; Duan, Q.
Show abstract
Biological pathway analysis often requires identifying interventions that block reachability to an undesirable state, such as a disease-associated module, toxic byproduct, or adverse phenotype, while preserving reachability among essential biological functions. Motivated by this setting, we study the Reachability Preserving Minimum Edge Cut (RPMEC) problem: given protected terminals s1 and s2 and a target terminal t, the goal is to remove a minimum-cost set of edges that separates s1 and s2 from t while keeping s1 and s2 connected. This formulation naturally models pathway-level intervention design, where one seeks to disrupt harmful signaling, metabolic, or interaction routes without breaking required functional connectivity. We revisit the three-terminal undirected edge-cut case and analyze a Dijkstra-style dynamic programming algorithm that is exact on planar graphs but fails on general graphs. We characterize the structural condition required for exactness, namely frontier-realizability of optimal source-side regions, and identify biological graph representations where this condition is likely to hold after appropriate preprocessing, including curated planar pathway maps, Reactome-style hierarchy trees, SCC-contracted feedback modules, metabolic building-block DAGs with dominator structure, and functional-module quotients of protein interaction networks. We further discuss directed variants, approximation strategies, and exact alternatives based on ASP, MILP, bounded-treewidth dynamic programming, and important separators. The results provide a graph-theoretic foundation for deciding when fast greedy computation is reliable for biological pathway intervention problems and when more expressive exact optimization methods are needed. Author SummaryMany real-world networks require interventions that separate harmful or undesirable states while preserving essential connectivity. This situation appears in biological pathway analysis, where one may want to block reachability to a disease-related module, toxic byproduct, or adverse phenotype without disrupting communication among essential genes, proteins, reactions, or metabolites. We study this problem through the Reachability Preserving Minimum Edge Cut formulation. Unlike ordinary minimum cut, the solution must satisfy both a separation requirement and a preservation requirement. We show why a natural Dijkstra-style algorithm works only under specific structural conditions, such as planar, laminar, or module-like pathway graphs, and why it may fail on general graphs. The results help identify when fast graph algorithms are reliable for biological intervention problems and when exact optimization tools such as Answer Set Programming or integer programming are more appropriate.
YADAV, P.; Singh, A.
Show abstract
The brain is the most captivating chef doeuvre of nature. Naturally then, the mind wonders about the process that births such a fascinating organ. Neurodevelopment is a complex yet robust phenomenon that conceals answers to our questions in its intricacies. In an attempt to shed some light on this matter, we study the developing brain connectome of the nematode, C. elegans across the post-embryonic phase. A tiny organism with only around 200 neurons comprising its brain and yet a diverse array of behaviors to display, it makes for a great model. Starting with most of its head neurons already present at hatching, the worm brain accumulates numerous more synaptic connections increasing the edge density. It maintains a weak connectivity throughout thereby, balancing global communication as well as hierarchy. At the mesoscopic level, we find that the core has a conserved backbone of persistent neurons along with a dynamic component formed of transient/recurring neurons. Moreover, the connectome has a rich club organization since the early stage which selectively strengthens indicating progressively denser connectivity among the integrators due to the previously reported asymmetric synapse addition. This asymmetry also shows up in the preservation of input hubs across development and the progressively more centralized organization of the in-degree k-core. Our work provides a new perspective into the neurodevelopment of the brain that may facilitate our understanding of its functioning.
Cahill, K. J.; Dhamala, M.
Show abstract
Understanding how complex systems self-organize, exhibit emergent properties beyond their constituent elements remains a challenge across physics, biology, and cognitive science. In resource-constrained neuronal systems, existing theoretical approaches, including gauge theoretic formulations, statistical physics-inspired methods, dynamical population models, and variational principles such as the Free Energy Principle, address important aspects of this problem but do not fully specify the physical conditions and thermodynamic costs under which self-organizing behavior occurs. Here, we introduce Dynamic Resource Theory (DRT) as a general physical framework for describing self-organization under constrained resource availability. DRT formalizes complexity as a physical property of self-organizing systems arising from coupled mechanisms of resource allocation and dynamic reallocation of internal resources. This framework provides a thermodynamic and variational account of how stability is preserved while adaptive reconfiguration remains possible, consistent with stationary action and thermodynamic constraints. DRT is formulated within a gauge theoretic setting and directly incorporates the energetic costs associated with maintaining structure and enabling system-level reconfiguration. Within DRT, baseline resource allocation preserves system stability, while internal and external demands perturb the system, driving self-organization through dynamic resource reallocation across a coupled free energy landscape without assuming subsystem separability. We then develop Neural Resource Theory (NRT) and Cognitive Resource Theory (CRT) as principled specializations of DRT, illustrating how this structure is instantiated in resource constrained neuronal and cognitive systems. We conclude by discussing the broader implications of DRT for understanding how complexity, emergence, and adaptive capacity arise over time through thermodynamically permissible reallocation processes across scales.
Michels, J. J.
Show abstract
Biomolecular condensates that form via liquid-liquid phase separation (LLPS) of, most prominently, intrinsically disordered proteins (IDPs) are ubiquitous in eukaryotic cells and responsible for regulating a plethora of biological functions. Amongst these, they contribute to regulating cell motility, either individually within an extracellular matrix or collectively within confluent epithelial tissue. In this computational study we focus on the latter with the aim of investigating whether the mutual exertion of mechanical forces during collective migration in an epithelium can principally trigger cytoplasmatic LLPS. Since present models for confluent epithelial motility have so far only considered cells that are devoid of phase separating (protein) solutes, we extend a common multiphase approach for 2D cell motility with a mixing contribution including any number of protein solutes. Our model considers the phase behavior in both intracellular and extracellular regions and determines to what extend the membrane is permeated by the solutes under the influence of mechanical and osmotic forces. Our initial calculations unlock a very rich behavior involving formation and dissolution of condensates during migration, as well as an impact of LLPS on the very nature of the motility itself, through feedback mechanisms which may bear biological relevance.
Deng, J.; Zhang, X.; Zhang, X.; Yang, X.
Show abstract
Coupled diffusion-reaction partial differential equations (PDEs) describe biochemical network dynamics but are difficult to solve for realistic multi-species systems without combining mechanism and data. We present a multi-stage physics-informed neural network (PINN) for multi-species diffusion-reaction PDEs and apply it to two ordinary-differential-equation (ODE) reference systems: the Boehm et al. JAK-STAT5 signaling pathway and the Sturis ultradian insulin-glucose model. For STAT5 we pose a latent-species identifiability test: given sparse observations of eight species, a ten-species model that retains two deliberately withheld but mechanistically standard components--an active receptor-JAK complex and the SOCS negative-feedback inhibitor--recovers the reference trajectory and reduces mean root-mean-square error 3.1-fold relative to an eight-species model that omits them, whereas a PDE-only solution without data anchoring diverges. Because the reference is itself ODE-generated, this demonstrates identifiability against synthetic data, not the discovery of new biology. For the insulin-glucose model the same framework reproduces the [~]120-minute oscillation to 1.0% mean relative error as a benchmark on a stiff, multi-timescale oscillator; its spatial dimension is treated as a numerical construct, not a physical transport setting. A Lyapunov analysis of the STAT5 ODE returns a maximal exponent statistically indistinguishable from zero ({lambda}max {approx} 3.61 x 10-5 min-1, 5/8 trials positive; Lyapunov time [~]1.9 x 104 min, far exceeding the 240-720 min horizon), so the system is effectively non-chaotic and the relevant instability is a bounded, parameter-induced trajectory divergence. Anchoring the solution to baseline data suppresses this divergence, with the reduction growing monotonically with sampling density--from [~]15-19% at eight time points to [~]88-97% at sixty-four, depending on perturbation magnitude. The framework thus offers a data-anchored route to latent-species identifiability and divergence suppression in biochemical ODE/PDE systems, demonstrated here against synthetic reference data. Inside cells, a three-dimensional chemistry of diffusing, reacting molecules drives signaling and rhythm--dynamics that, for realistic networks, strain conventional solvers. Here a multi-stage physics-informed neural network--machine learning constrained by the governing equations--solves stiff, multi-species reaction systems from sparse data. In the JAK-STAT5 signaling pathway, a model that retains two standard but unobserved components (an active receptor complex and a negative-feedback brake) recovers a reference trajectory that a reduced model cannot--a controlled test of whether sparse data can pin down withheld pecies, not a claim of new biology. The same framework reproduces the roughly two-hour insulin-glucose rhythm to within 1% as a benchmark on a stiff oscillator. And anchoring the solution to a few dozen baseline measurements collapses parameter-induced trajectory divergence, turning a parametrically sensitive simulation into a stable one. Where mechanism and data meet, sparse measurements can constrain the structure a model would otherwise leave undetermined.
Charan, K.; Kar, S.
Show abstract
In mammalian cells, under normal circumstances, the p53 protein exhibits oscillatory dynamics in response to DNA damage and maintains the cells in a cell-cycle-arrested state. Intriguingly, some cells can escape this cell-cycle-arrested state even after prolonged DNA damage, and often undergo mitotic catastrophe. In this context, the precise role of p53 dynamics and its complex interplay with cell-cycle regulation remain poorly understood. Herein, by constructing a comprehensive network model, we have identified crucial crosstalk regulations between the p53 protein and key cell-cycle regulators that enable some cells to escape cell-cycle arrest during prolonged DNA damage. The model further illustrates a probable cellular mechanism underlying mitotic catastrophe and predicts ways to induce it in a therapeutically relevant manner.
Senguler Ciftci, F.; Erman, B.
Show abstract
This study introduces a statistical mechanical framework for allosteric communication in proteins based on the spanning-tree ensemble of residue contact networks. By representing protein structures as weighted graphs, we identify each spanning tree as a topological microstate. The canonical partition function is evaluated exactly via the determinant of the reduced weighted Kirchhoff (Laplacian) matrix, allowing for the derivation of global thermodynamic functions (including Helmholtz free energy, internal energy, entropy, and heat capacity) without approximation. Allosteric channels between specific residue pairs are defined as sub-ensembles containing unique simple paths. Using the Burton-Pemantle theorem and the Moore-Penrose pseudoinverse of the graph Laplacian, we compute exact path probabilities and channel-specific thermodynamics. This methodology enables a decomposition of channel heat capacity into energetic and topological components and quantifies residue-level allosteric importance through fractional contributions to the channel partition function. The framework was applied to the G12D mutation in KRAS, comparing wild-type (PDB: 6GOD) and mutant (PDB: 6GOF) proteins. Results show that while the mutation minimally affects mean internal energy and entropy, it reduces global heat capacity by 27.3%. This indicates a topological stiffening where the mutant occupies a significantly narrower landscape of spanning-tree configurations. At the channel level, the mutation maintains distributional stability across six functional routes but triggers a substantial internal redistribution of allosteric importance. Specific residues, such as Q61 and F156, shift occupancy by up to 35.5%. These findings suggest that the G12D mutation does not destroy communication pathways but reorganizes internal information traffic to favor a catalytically impaired state. This approach provides a rigorous, parameter-free metric for understanding how point mutations perturb distal protein signaling.
Alibutud, R. F.; Kumar, S.
Show abstract
Phylogenetic inference is a common task in molecular and evolutionary biology and has conventionally required a multiple sequence alignment (MSA), a statistical model of amino acid substitutions, and an optimality principle. Recently, global models of amino acid substitutions have been inferred from millions of MSAs using transformer-based deep learning, resulting in protein foundation models (pFMs), also known as protein language models (PLMs). Training pFMs on MSAs hypothetically enables them to encode residue dependencies and the phylogenetic structure of the MSA collection. In contrast, pFMs trained on individual sequences lack access to such phylogenetic structure. Here, we assess the phylogeny inference gains offered by the use of MSA for training pFMs by comparing the relative accuracies of phylogenies inferred using two types of pFMs: one trained on a large collection of MSAs (msat-pFM, [1]) and the other trained using a collection of single sequences (esm-pFM). For msat-pFM analysis, we inferred neighbor-joining trees using pairwise distances estimated directly from the sequence attention matrices. For esm-pFM [2], pairwise distances were obtained using the correlation of attentions of homologous residues, where pairwise sequence alignments (PSA) were used to establish residue homologies. Surprisingly, MSA phylogenies inferred using the msat-pFM were less accurate than esm-pFMs. This pattern was seen across datasets spanning both small and large numbers of species and proteins. Also, PSA phylogenies obtained using residue attentions from early ESM-PFM layers were much more accurate. These results suggest that the multiple sequence alignment step, which is obligatory to establish residue homologies across multiple sequences, may not add information when using evolutionary distances based on attentions in pFMs.
Zhao, D.; Yang, Y.; Sun, J.; Zhang, J.; Duan, H.; Tan, Y.; Liu, l.
Show abstract
Although the "RNA world" hypothesis suggests that RNA played a crucial role in the origin of life [7], the functional framework of RNA in prebiotic protein synthesis and the mechanisms of genetic code formation during the prebiotic period remain poorly understood. Here, using the prebiotic "primordial soup" as a model, we reconstructed the detailed steps that would yield a protein with a stable ordered amino-acid sequence in the "primordial soup" at the prebiotic period. In the "primordial soup", a large number of medium- to large-sized biomolecule-like substances--such as RNA-like and protein-like molecules of various sizes and shapes, as well as related polymers like amino-acid-RNA-like etc.--did generate and accumulate. Moreover, protein-like and RNA-like molecules formed even more intricate complexes. These complexes bound free mRNA-like molecules through complementary base pairing. Subsequently, with an extremely low probability, two adjacent amino-acid-RNA-like molecules became bound to this free mRNA-like molecule, and their amino acids underwent a condensation reaction by the complexes, producing peptides and eventually proteins or polypeptides. This free mRNA-like molecule exhibits a certain flexible structure, whereas the super-large complexes formed by protein-like and RNA-like molecules (which possess certain activities) and the amino-acid-RNA molecules exhibit relatively rigid structures. Long-term evolution and mutual selection led to the emergence of proteins with stable amino acid sequences and moderate catalytic activity. In this way, the nucleotide information embedded in such mRNA-like molecules indirectly express through protein synthesis--a process we term the "A Co-Adaptation Flexible-Rigid Docking Model", where flexible mRNA-like molecules dock onto rigid complexes to enable ordered peptide formation. Finally, we show how trinucleotide codons emerge naturally from the flexible-rigid docking constraints.