Back

Assessing nonlinearities in the GFP random mutagenesis landscape using the Power Transform

Petrov, D. A.; Ivankov, D. N.

2025-12-12 bioinformatics
10.64898/2025.12.09.693289 bioRxiv
Show abstract

Epistasis, a non-additive contribution of mutations to fitness, complicates genotype-to-phenotype prediction and is often confounded by nonlinearities in phenotype measurements. The Power Transform method, particularly Box-Cox, has been used to reduce epistasis in small, combinatorially complete datasets by linearizing phenotypic scales. However, its applicability to large landscapes, especially those generated by (quasi-)random mutagenesis, remains unclear. Here, we apply both Box-Cox and Yeo-Johnson Power Transforms to the extensively characterized green fluorescent protein (GFP) fitness landscape, which contains hundreds of thousands of two- and three-dimensional combinatorially complete datasets. Surprisingly, when applied to individual hypercubes, Box-Cox reduces pairwise epistasis by 19.4% on average, whereas Yeo-Johnson - despite accommodating negative values - slightly increases epistasis (by 4.45%). Moreover, in a notable fraction of hypercubes, both methods increase the magnitude of epistatic coefficients, indicating that Power Transform does not universally reduce nonlinearity at the local scale. In contrast, applying Yeo-Johnson to the largest connected component of the GFP landscape (20,872 genotypes) successfully reduces both pairwise (by 3.86%) and third-order epistasis (by 11.6%), with the majority of coefficients decreasing in magnitude. These results demonstrate that Power Transform can be extended to large, real-world landscapes generated by random mutagenesis, but only when applied to connected sublandscapes. Our findings highlight a critical distinction between global and local linearization and caution against assuming that Power Transform always diminishes epistasis.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.