Entropy
○ MDPI AG
Preprints posted in the last 90 days, ranked by how well they match Entropy's content profile, based on 21 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Malatesta, P.; Chandnani, R. S.; Yalim, J.; Ozkan, S. B.; Panagiotou, E.
Show abstract
MotivationWith the rapid development of AI methods that predict protein structures from sequence, understanding the structure-function relation increasingly depends on quantitative structural descriptors that are both biologically meaningful and scalable to large datasets. Here, we introduce mathematical topology metrics that quantify the entanglement complexity of a tertiary protein structure while respecting uncrossability constraints. ResultsBy employing only three such metrics across all protein structures in the Protein Data Bank, we represent the proteome structural space in a continuous three-dimensional space. Distances within this space capture structural similarity and correlate with functional similarity. We find that the mathematical entanglement based landscape of protein structural space diversifies with the evolutionary expansion of protein function across species. Moreover, this continuous representation reproduces CATH classifications with high accuracy for major structural classes. These results indicate that these metrics efficiently encode structural features linked to protein function and provide a more informative description than conventional metrics. AvailabilityData used in this study are available in the Protein Data Bank. Details of the machine learning model used can be found in https://github.com/roshitac/CATH_Classification-. ContactBanu.Ozkan@asu.edu, Eleni.Panagiotou@asu.edu Supplementary informationSupplementary data are available at Journal Name online.
Huang, Q.; Guo, H.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWCellular automata and graph reaction-diffusion systems encode local spatial interactions in different mathematical forms. We develop a cochain-operator calculus for these two settings. Over a finite field Fq, every local rule on a finite neighborhood has a unique reduced polynomial representative. On an oriented line, the coboundary and endpoint maps recover the left and right shifts. Our main theorem shows that these operators, together with linear operations, constant cochains, and the degree-zero cup product, generate every finite-radius polynomial cellular automaton. Explicit formulas for Rules 30, 110, and 22 show how reflection-invariant linear coupling, directed transport, and nonlinear neighbor interactions enter the calculus. On a general graph, d*d is the unweighted combinatorial Laplacian and enters a graph reaction- diffusion recurrence. Over [R], the term - Dd*d with D [≥] 0 admits the usual diffusion interpretation; over Fq, the corresponding expression defines modular coupling without an intrinsic order. In the morphogenetic examples, we therefore distinguish pattern-generating dynamics from finite-state observation and use the Betti numbers of active induced subcomplexes to summarize observed patterns. This yields a common algebraic representation without identifying real-valued diffusion with finite-field dynamics.
Kringelbach, M. L.; Deco, G.
Show abstract
Brain dynamics can be described in three different convenient mathematical languages, namely connectome harmonics, turbulence and complex harmonics (CHARM). Here we demonstrate that these theoretical frameworks can be rigorously unified, under the functional calculus, as one self-adjoint operator and its single spectral measure. The connectome Laplacian carries that measure; the harmonics are its spectral projections, the turbulence smoothing kernel is its resolvent, and the CHARM form is its unitary propagator. The bridge that makes this exact is a textbook fact: The exponential distance rule, which is the empirical kernel of the turbulence model, is the Greens function of a screened Laplacian, so the local order parameter is the phase field passed through the resolvent of the same operator whose eigenfunctions are the harmonics. A single shared control parameter, the spectral gap, simultaneously yields the cortical hierarchy, the turbulent information cascade and the structured interference the CHARM form measures. This unification makes a strong predictive claim. If the harmonic projections, the turbulence resolvent and the CHARM propagator really are three functions of one operator, then any structural perturbation that re-tunes the operator must move all three signatures in unison and must do so with a single coupling. We test this prediction with a pharmacological perturbation by lysergic acid diethylamide (LSD), which is known to change the emotional state, by empirically perturbing the operator with a 5-HT2A receptor density map and asking whether one scalar coupling can simultaneously predict the multi-scale turbulence shift observed, through the resolvent, and the macroscale harmonic energy redistribution, through first-order Rayleigh-Schrodinger perturbation theory. We found that the two independent functional domains respond in unison to one structural perturbation of one operator. The identity is exact as operator calculus and its purchase on the brain depends on a single load-bearing seam, the degree heterogeneity of the connectome, which we make explicit. We propose that this single-operator structure is the necessary mathematical scaffolding of our Entangled Loop theory.
Ali, A. F.; Inan, N.; Laukkonen, R.; Mikheenko, P.
Show abstract
We develop a theoretical proposal linking vacuum stability and brain dynamics through superconductivity-inspired coherence, symmetry reduction, and the thermodynamic stabilization of low-entropy regimes. We take an unbroken SU(3) structure as a candidate stable residue of the low-temperature vacuum. At the neural level, we formulate a coarse-grained analog in which a two-fluid model with dissipative and coherence-supporting components describes brain dynamics. Specifically, the coherence-supporting component is proposed as a possible basis for the efficient binding and integration required to sustain a stable, unified conscious state. The proposal offers a common geometric language for relating physics and neuroscience with falsifiable signatures in coherence and state-dependent transitions. The main technical contribution is a computational algebraic model of conscious-state dynamics, where neural data are mapped to reconstructed state trajectories. Effective generators are inferred from those trajectories, and the two-fluid split is tested as a Cartan-root decomposition of su(3), with a rank-two commuting sector for coherence-preserving balance and six root directions for state transitions. This structure can be tested on neural data and contrasted with alternative dynamical models.
BV, H.; Adigwe, S.; Jolly, M. K.; Gedeon, T.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWCell fate decisions are driven by gene regulatory networks (GRNs). While the mutually inhibitory toggle switch effectively models binary fate decisions, fully connected inhibitory networks with more than two nodes fail to capture multi-fate decisions due to the low prevalence of "single high states", where only a single master regulator is highly expressed. The goal of this study is to find network structures that support all single high states. We find that the only network that attains the highest possible prevalence of all single high states within the set of monotone Boolean (MB) models is completely disconnected. Since biological networks typically require connectivity, we investigate network structures that support equipotency, where all single high states have equal prevalence within MB models. Finally, we characterize the networks that support multistability between all single high states, finding that it is possible only in networks in which each node either has self-activations or is inhibited by every other network node. Our findings provide a theoretical framework for understanding the network design principles that can support simultaneous differentiation into multiple distinct cell types.
Varshney, D.; Tajjar, M. H.; de Vries, J.; Hutter, F.; Rensing, S. A.
Show abstract
How morphological complexity evolves is still enigmatic. While there is evidence in algae and plants as well as animals that diversification of the repertoire of transcription factors (TF) is causative for evolution of organismal complexity, there are many examples from lineages that follow their own way of complexity evolution, for example by expansion of particular families. For land plants, correlation of the size of the TF complement with number of cell types (as a proxy for morphological complexity) has been shown, and several families were identified as candidates to drive complexity evolution. Here, we expand a previously available dataset of cell type numbers from 12 to 82 proteomes and introduce a four class body plan scheme. We find that the total TF complement correlates with the number of cell types of Archaeplastida (primary plastid bearing plants and algae). We used TabPFN (Tabular Prior-data Fitted Network) for binary (uni- vs. multicellularity) as well as for four class Bauplan classification. TabPFN is able to predict the morphological complexity with high accuracy. This approach allows to determine organismal complexity based on the gene space of an organism. Based on our results, we can confirm that plant morphological evolution is driven by gain and expansion of TF families.
YADAV, P.; Singh, A.
Show abstract
The brain is the most captivating chef doeuvre of nature. Naturally then, the mind wonders about the process that births such a fascinating organ. Neurodevelopment is a complex yet robust phenomenon that conceals answers to our questions in its intricacies. In an attempt to shed some light on this matter, we study the developing brain connectome of the nematode, C. elegans across the post-embryonic phase. A tiny organism with only around 200 neurons comprising its brain and yet a diverse array of behaviors to display, it makes for a great model. Starting with most of its head neurons already present at hatching, the worm brain accumulates numerous more synaptic connections increasing the edge density. It maintains a weak connectivity throughout thereby, balancing global communication as well as hierarchy. At the mesoscopic level, we find that the core has a conserved backbone of persistent neurons along with a dynamic component formed of transient/recurring neurons. Moreover, the connectome has a rich club organization since the early stage which selectively strengthens indicating progressively denser connectivity among the integrators due to the previously reported asymmetric synapse addition. This asymmetry also shows up in the preservation of input hubs across development and the progressively more centralized organization of the in-degree k-core. Our work provides a new perspective into the neurodevelopment of the brain that may facilitate our understanding of its functioning.
Li, D. J.
Show abstract
All cellular life forms fall under the three-domain classification of life, raising a fundamental evolutionary question: why does this classification feature three rather than two or four? To answer this question, a more general method, rather than the traditional one based on comparing small-subunit ribosomal RNAs, is required. The three-base periodicity in genomes is a common feature of both cellular life forms and viruses, which is species-specifically biased between amino acid biosynthetic families. Based on comparing such a common feature of all life forms, a global triangular diversification picture has been obtained, whose three angular regions correspond to the three domains, respectively. This mechanism of diversification of life attributes the evolutionary driving forces in diversification of the three domains of life to the biases between amino acid biosynthetic families. Notably, the same mechanism also applies to the contemporary diversification of SARS-CoV-2, whose reasonable results in turn corroborate the above explanation of primordial diversification of life and in addition shed light on the mechanism of speciation.
KUNDU, S.
Show abstract
Small molecule modifiers whence bound, allosterically, will alter the binding of a macromolecule to one- or more-cognate substrates/partners via conformational and non-conformational changes. Although allostery is inferred directly from empirical data, the mathematical basis of these models, constraints deployed and choice of parameter(s) are not clear. Here, we present and characterize a discrete-to-continuous mathematical model for ensemble distributions of a ligand-interacting macromolecular species across milieux-dependent conformational states and examine its role in the genesis and progression of cooperative binding. The premise, of our model, is a set of occupancy matrices (sparse, binary, strictly delocalized) which can be partitioned by a probability-based hyperparameter into mutually exclusive proper subsets of occupancy matrices with identical multinomial probabilities. Since each subset is canonical with a constituent occupancy matrix, it is characterized by a unique multinomial probability. The inner product of combinatorial pairs of all mutually exclusive subsets of occupancy matrices, with an expression for the summed transitional probabilities (finite differences between unique multinomial probabilities), is the differentiable matrix of strictly positive real-valued numbers for the system of ensemble distributions. Whilst the harmonic mean is presented as a generic solution for a system of ensemble distributions, the row-wise definite integral for each column is the finite union of open intervals (contiguous, strictly monotone) which in tandem with a set of interval-specific and bounded transitional probabilities constitutes a piecewise smooth curve (path-connected-, closed- and compact-set). Our discrete-to-continuous model is phenomenological and able to recapitulate the basic tenets of cooperative binding whilst offering insights into the genesis and progression of the same.
Qun, Z.; Huaizheng, Z.; Yuxin, Z.; Jieying, B.; Tan, S.
Show abstract
Network centrality is the workhorse of gene prioritisation, yet what a ranking omits is rarely audited. Scoring each selection against an annotation-count-matched maximum-entropy reference--asking whether a selected gene set covers the genomes functional space or collapses it-reveals that the criterion in standard use has a measurable blind spot in exactly the class it is meant to surface. Degree, the most widely used criterion, returns the cross-module bridges that are also locally dominant--connector hubs--and omits the non-hub connectors: where 26% of the genome occupies these coordinating roles, a degree-ranked list holds 18% and an EDVS-ranked list 55%, and degrees top-1% collapses functional coverage below the reference on all five networks tested. We repurpose EDVS (Entropy of Degree-Vector Sums), an information-theoretic diversity measure, as an annotation-free, partition-free centrality that recovers this omitted class. The coverage it preserves is carried by cross-module participation P, which cannot be computed without a community partition; EDVS matches P-level coverage on all five networks using none, and retains 0.84 of its selection under edge perturbation that leaves partition-based selections at 0.21-0.46. The deficit is general: the collapse holds in the same direction on the two networks built without functional annotation (0.5-1.1 bit; co-expression, physical interaction) as on the three supervised by it (1.6-3.3 bit; RiceNet, AraNet, STRING), so supervision amplifies it rather than creates it. The remedy is bounded: EDVS ceases to preserve coverage on the sparse physical-interaction network. And the class EDVS isolates is organizational, not an importance signal: pre-registered probes--essentiality, transcription-factor identity, tissue-specificity, date/party-hub character, phenotype co-localisation--return null or reversed throughout. The conclusive ones are equivalent to their degree-matched nulls within {+/-}5 percentage points (demonstrated, not merely undetected), and the classical coupling of centrality to importance itself holds only network-dependently. Author SummaryGenes rarely act alone: many diseases and agricultural traits are shaped by genes that coordinate several biological processes rather than specialising in one. The standard way to find such genes in a network of gene interactions is to count each genes connections--its "centrality"--and rank genes by that count. We show this standard approach has a blind spot: it favours genes that dominate one process over genes that quietly bridge several processes without dominating any, and this blind spot appears across rice, thale cress, and yeast gene networks. We repurpose a diversity measure from an unrelated field (originally used to compare citation patterns) as a new way to rank genes that finds these bridging genes from network structure alone, without needing gene-function annotations--which are themselves incomplete and biased toward well-studied genes--or a prior, unstable step of splitting the network into modules. We are careful to show where the new approach also falls short: on sparse, noisy networks it stops working, and the genes it recovers are not shown to be more biologically important than other genes, only differently positioned. What that position is for is a question this work leaves open.
Izuazu, C.; Browne, C.
Show abstract
Mathematical models, e.g. differential equations and stochastic processes, have gained considerable attention for understanding evolution of antibiotic resistance. However, most existing models assume standing genetic variation and do not consider the possibility of random or drug-induced mutation of reference bacterial strains. Therefore, we propose a pharmacokinetics/pharmacodynamics (PK/PD)-based continuous-time Markov chain considering the competition and mutation between sensitive and resistant bacterial within an infected host during treatment. The proposed model is approximated as a generalized birth-death process with immigration, allowing for explicit derivation of the probability resistant population establishes during treatment. Besides capturing the stochasticity of de novo emergence of a resistant bacterial strain, we explore the effects of different antibiotic modes of action, horizontal gene transfer, nutrient availability and drug pharmacokinetics on antibiotic resistance. We find that replication-targeting (biostatic) drugs suppress resistance more than death-targeting (biocidal) drugs. Like prior works, we obtain maximized resistance at intermediate drug concentrations, however the consideration of de novo mutation magnifies the superiority of higher doses in preventing resistance emergence.
Zhao, D.; Yang, Y.; Sun, J.; Zhang, J.; Duan, H.; Tan, Y.; Liu, l.
Show abstract
Although the "RNA world" hypothesis suggests that RNA played a crucial role in the origin of life [7], the functional framework of RNA in prebiotic protein synthesis and the mechanisms of genetic code formation during the prebiotic period remain poorly understood. Here, using the prebiotic "primordial soup" as a model, we reconstructed the detailed steps that would yield a protein with a stable ordered amino-acid sequence in the "primordial soup" at the prebiotic period. In the "primordial soup", a large number of medium- to large-sized biomolecule-like substances--such as RNA-like and protein-like molecules of various sizes and shapes, as well as related polymers like amino-acid-RNA-like etc.--did generate and accumulate. Moreover, protein-like and RNA-like molecules formed even more intricate complexes. These complexes bound free mRNA-like molecules through complementary base pairing. Subsequently, with an extremely low probability, two adjacent amino-acid-RNA-like molecules became bound to this free mRNA-like molecule, and their amino acids underwent a condensation reaction by the complexes, producing peptides and eventually proteins or polypeptides. This free mRNA-like molecule exhibits a certain flexible structure, whereas the super-large complexes formed by protein-like and RNA-like molecules (which possess certain activities) and the amino-acid-RNA molecules exhibit relatively rigid structures. Long-term evolution and mutual selection led to the emergence of proteins with stable amino acid sequences and moderate catalytic activity. In this way, the nucleotide information embedded in such mRNA-like molecules indirectly express through protein synthesis--a process we term the "A Co-Adaptation Flexible-Rigid Docking Model", where flexible mRNA-like molecules dock onto rigid complexes to enable ordered peptide formation. Finally, we show how trinucleotide codons emerge naturally from the flexible-rigid docking constraints.
Tampakaki, A. E.; Barmparis, G. D.; Angelaki, E.; Marketou, M. E.; Tsironis, G. P.
Show abstract
We present a quantum-enhanced version of the classic k-Nearest Neighbors (kNN) classification algorithm, applied to the prediction of arterial hypertension. The traditional Euclidean distance metric of the kNN algorithm is replaced with a Fidelity-derived quantum dissimilarity measure to evaluate the similarity between data samples. We map classical real-world clinical and ECG-derived data features into quantum states via the Dense-Angle Encoding, which efficiently utilizes parameterized rotation gates to pack multiple features into minimal qubits while maintaining pure states. We evaluate the performance of the dissimilarity measure using both the noiseless state vector Simulator and the IBM Qiskit Estimator primitives. The quantum circuit demonstrates robust predictive capabilities comparable to the classical model. While it does not claim computational supremacy over the classical baseline, the framework proves that fidelity-based similarity is a physically meaningful and efficient approach for hybrid quantum classical classification.
Longhi, C.; Martinez-Vaquero, L. A.; Trianni, V.
Show abstract
Many proposed mechanisms for the evolution of cooperation among unrelated individuals rely on relatively demanding cognitive abilities that are not widespread across taxa. In contrast, individual heterogeneity is a pervasive feature of animal groups, encompassing differences in personality as well as physical and cognitive traits. Such heterogeneity can promote the evolution of cooperation, yet its role has received comparatively little attention, particularly as a source of variation giving rise to social organization such as leadership. A specific form of leadership can emerge under unstable environmental conditions, when some individuals become better suited than others to initiate action and influence the behavior of their peers. Unlike fixed dominance hierarchies, emergent leadership can rapidly adjust to changing environmental conditions, thereby reshaping group organization. Because it does not require the maintenance of stable hierarchies, this form of leadership can arise even in species that do not have the cognitive capabilities to sustain complex social structures. In this work, we investigate the combined effects of individual heterogeneity and emergent leadership on the evolution of cooperation using an evolutionary game-theoretic model in which individuals may assume the roles of leaders or followers according to their strength, representing individual differences in suitability to prevailing environmental conditions. We examine different levels of population heterogeneity together with increasingly complex strategy sets requiring progressively greater informational requirements, allowing individuals to condition cooperation on their own strength, leadership role, or both. Our results show that the interplay between leadership and heterogeneity promotes the evolution of cooperation, particularly when only a small fraction of individuals act as leaders. Under these circumstances, cooperation evolves even when individuals employ the simplest possible strategies. Under harsher ecological conditions, cooperation can be sustained by more sophisticated strategies, specifically by conditional strategies that prescribe cooperation when individuals are strong or leading and defect when acting independently. Author summaryIn this study, we propose that emergent leadership mediated by individual diversity can boost the evolution of cooperation in animal groups. Building on growing evidence on the heterogeneity of animal capabilities and personalities, we focus on the fleeting leadership that emerges in animal groups when facing rapidly changing environmental conditions. We suggest that this type of leadership that emerges from individual differences in strength--a generic quality encompassing those characteristics that make an individual more fit to lead in a given situation--does not require complex cognitive capabilities from the animals and represents a valid alternative to more demanding strategies proposed in the past to explain the evolution of cooperation. Using an evolutionary game theory model, we show that if a population includes a few strong players, these can become influential leaders and guide the actions of their peers to achieve cooperation. Although the naive strategy of always cooperating is sufficient for cooperation to evolve, the introduction of more complex strategies leads players to cooperate only when they are more likely to be recognized as influential leaders. These strategies are more effective in promoting cooperation under unfavorable ecological conditions and are also more robust against exploitation by defectors.
Margarit, D.
Show abstract
Structural network representations of metastatic dissemination typically focus on static topology without resolving transport dynamics, relaxation timescales, or steady-state behaviour. Here, we formulate a discrete Markovian transport model on a directed higher-order network with transition rates derived from qualitative clinical affinity classes. By constructing a non-Hermitian row-stochastic transfer operator, we characterise the relaxation dynamics through its spectral decomposition. The system exhibits a fast-mixing regime characterised by a spectral gap of {gamma} {approx} 0.67, corresponding to a characteristic relaxation timescale of {tau} {approx} 1.49 discrete steps, with the influence of the primary tumour origin progressively attenuated during dissemination. Convergence towards a non-equilibrium steady state (NESS) is accompanied by a reduction in Shannon entropy, concentrating probability mass within specific topological sinks. This spectral relaxation delineates two distinct dynamical regimes: early transient dissemination (n < {tau}), dominated by local organ-specific transition probabilities (organotropism), and the asymptotic regime (n > {tau}), determined increasingly by the global transport architecture of the network. Comparison with independent clinical and autopsy observations across 21 primary tumours and 23 target organs indicates that the predicted stationary distribution is consistent with the observed hierarchy of metastatic organ involvement.
Sadhukhan, S.; Santra, D.
Show abstract
Diffuse gliomas are deadly because the individual tumor cells invade - they travel far from the imageable mass, so it is impossible to remove the tumor completely. On the cellular level, glioma cells seem to be in either a "go" state (in which they do not divide) or a "grow" state (in which they do not migrate). We investigate what this tiny choice has to say about the large-scale speed of the invasion front and whether the implication is sufficiently strong to rule out the classical description of the Fisher-Kolmogorov-Petrovsky-Piskunov (Fisher-KPP) type, in which a single phenotype migrates and proliferates. We derive a two-phenotype reaction-diffusion model with density-dependent switching, and we prove the cooperative (quasi-monotone) structure and the associated comparison principle and study travelling-wave solutions of the model. A leading-edge linearization gives minimal front speed as minimizer of an explicit dispersion relation, and direct simulation verifies the predicted speed. In the experimentally relevant fast switching limit, we find a closed-form expression for the speed, that is, we obtain an effective Fisher-KPP equation with rescaled diffusivity and growth rate, with the fractions of the phenotypes. The "go-or-grow" (GoG) front can move at a maximum speed of half the Fisher speed for the same single-cell motility $D$ and proliferation rate $r$, which occurs only when the cells divide their time equally between the two phenotypes. This bound is directly testable: measurement of the front speed, plus independent determination of $D$ and $r$, discriminates the two hypotheses, and in the GoG case, yields recovery of the phenotype balance. We then extend the result to anisotropic (DTI-informed) invasion along white-matter tracts and discuss implications for understanding clinical measurements of growth rate.
Desai, R.; Pople, D.; Musale, A.; Jain, S.; Sajjad, I.; Wittebort, R. J.; Koder, R. L.; Nanda, V.
Show abstract
The folding thermodynamics of proteins are dominated by two opposing forces, the loss in backbone entropy and the packing of hydrophobic groups. The same forces are major contributors to the extension thermodynamics of elastic proteins with the distinction that both processes act in concert, favoring the higher chain and solvent entropy of a relaxed conformation. The relative entropic contributions specify the recoil mechanism; human elastin recoil is primarily driven by hydrophobic forces, whereas fly resilin has a rubber-like mechanism driven by backbone entropy. Despite the importance of elastic proteins to tissue biomechanics, few have been identified, let alone characterized to the same extent as elastin and resilin. We develop a thermodynamic framework that maps proteins by sequence-derived estimates of extension-induced backbone and solvent entropy changes. Putative elastic proteins are proposed and classified by recoil mechanism based on estimated thermodynamic features. Proteins that map to elastic regions are overrepresented by the skin proteome. The set of predicted elastic domains is further extended by incorporating sequence context embedded in protein language models. Protein domains with distinct thermodynamic recoil mechanisms cluster on the latent space manifold. Some of these domains are anticipated to have roles within molecular machines, expanding the scope of elastic protein function beyond mechanical materials like elastin and resilin.
Zakar-Polyak, E.; Kerepesi, C.
Show abstract
Contextualized protein-protein interaction networks provide crucial insight into diseases and other biological processes, but for a profound understanding of such processes and their distinct effects on individuals, the protein-protein interactions within individual samples must be investigated. A straightforward approach to estimate the PPI network of a sample is to restrict a general network of known PPIs to the proteins that are found in the sample. Although proteomics methods are becoming more accessible and precise, large-scale and single-cell studies still mainly target characterizing the transcriptomics profile of the samples, which is then often used as an approximation of the protein activities. The correlation of gene expression and protein abundance has been addressed in the past, but information about the deviations of the different omics-based estimates of the PPI networks is still lacking. In this study, we performed a comparative analysis of transcriptomic-based and proteomic-based sample-specific PPI network estimates to fill this gap. We created a framework for a comprehensive and transparent comparison of the two omics levels in two independent datasets, with a special focus on time-related network dynamics. We found that the size-adjusted characteristics of the different omics-based networks are very similar; the overall trend of how they change with time is also often the same, but the rate of the changes typically differs. The characteristics of the nodes present in both types of networks also show high similarity and often different time-related rates of change, but this varies among metrics. These results shed light on the properties of PPI network estimations and advise caution in interpreting them appropriately.
Rabier, C.-E.; Berry, V.; Glaszmann, J.-C.
Show abstract
Asian rice is one of the best documented crops in terms of genetic diversity. The domestication process, that probably started 9000 years ago in China, remains difficult to infer since the main vertical signal is blurred by horizontal signals related to gene flow among cultivars and wild relatives. Consequently, a large number of hypotheses on the domestication process of rice have been published. Besides, most of the methods used to infer these scenarios do not model all the known biological phenomena at stake. Here, we present a methodological study based on a rich stochastic model, that incorporates introgression events, incomplete lineage sorting, and mutations that happen over time. The global evolutionary scenario is represented by a phylogenetic network. Furthermore, each locus scenario is modeled according to a locus tree through the Multispecies Network Coalescent. More importantly, for inferring the phylogenetic network, we propose a new hybrid approach combining a phylogenetic network method and a machine learning technique. In particular, our hybrid approach, named SO_SCPLOWNARFC_SCPLOW, benefits from advantages of a mathematical phylogenetic method, SO_SCPLOWNAPPC_SCPLOWNO_SCPLOWETC_SCPLOW, and from the potential of a powerful machine learning classifier, i.e. Approximate Bayesian Computation Random Forest (ABC-RF). These two methods are complementary since SO_SCPLOWNAPPC_SCPLOWNO_SCPLOWETC_SCPLOW reconstructs network accurately, whereas ABC-RF is able to handle a large amount of data. The originality is twofold. First, prior distributions required for ABC-RF are calibrated thanks to SO_SCPLOWNAPPC_SCPLOWNO_SCPLOWETC_SCPLOWs estimates. Secondly, ABC-RF relies on summary statistics inspired by phylogenetic network literature. We show, on simulated data, that the SO_SCPLOWNARFC_SCPLOW hybrid approach enjoys very good performances. On rice real data, it infers a scenario with a unique domestication (that of Japonica), followed by three reticulation events involving early Japonica. It highlights two introgression events at the origin of Indica and cAus, and one admixture event responsible for the emergence of cBas. Author summaryToday, in genomics, there is a real need for methods able to infer phylogenetic networks. A phylogenetic network is a directed graph representing events like hybridization, introgression, and horizontal gene transfer. Understanding these complex biological phenomena, essential for crop adaptation, can help breeders when facing challenges like climate change and population growth. Genome-wide diversity analysis thus requires network methods scaling for large data volumes and incorporating fundamental biological phenomena. In this context, we present a new hybrid approach, SO_SCPLOWNARFC_SCPLOW, that benefits from the potential of a powerful machine learning classifier, Approximate Bayesian Computation Random Forest, and from advantages of a mathematical phylogenetic method, SO_SCPLOWNAPPC_SCPLOWNO_SCPLOWETC_SCPLOW. Consequently, SO_SCPLOWNARFC_SCPLOW is able to handle large data-sets thanks to machine learning and is also based on a deep mathematical theory. On simulated data, our hybrid approach performs very well. When applied to real rice genomic data, it supports a scenario with a single domestication event, that of Japonica. The analysis further highlights the role of early Japonica in the origin of both Indica and circumAus. Finally, it identifies an ancient admixture event, involving circumAus in the emergence of circumBasmati. Together, these findings confirm the importance of early rice history along the Himalayan region.
Kondratev, A. Y.; Ianovski, E.; Voronina, E.; Crossa, J.
Show abstract
Multi-environment trials are central to cultivar evaluation because they reveal how candidate cultivars perform across locations, years, management conditions, and stress environments. The resulting yield matrix is a rich source of data on genotype-by-environment interaction, and a wide literature on estimation, decomposition, visualisation, and prediction of yield potential and stability has flourished. However the ultimate question of which cultivar to recommend on the basis of such a matrix is often left implicit. The question is far from trivial, and in this paper we formulate cultivar recommendation as an axiomatic ranking problem. This framework is rich enough to encompass the existing literature on stability indices, as well as any other deterministic ranking procedure. We show that many commonly used stability-based procedures can violate minimal criteria of efficiency or consistency. The result of such violations is that a cultivar with uniformly high yield could be ranked below a cultivar with uniformly low yield, or the relative ranks of two cultivars could depend on whether or not a third cultivar is present in the matrix. Our results prove that under a small number of such criteria the space of admissible rules collapses to the family of power means and their limiting cases. If we further wish to allow multiplication normalisation of yield, we are left with the geometric mean as the unique solution.