Back

Mutual information of high-dimensional random variables: estimation by frontier mutual information

Mori, T.; Kawamura, T.

2024-12-29 bioinformatics
10.1101/2024.12.29.630641 bioRxiv
Show abstract

Many natural sources of information, including genes, yield intricate datasets characterized by high-dimensional random variables. However, the computational constraints and information loss have often limited the accuracy of mutual information (MI) computations in such datasets. To address these limitations, we introduce a novel metric, micromutual information, which measures the information exchange at each cell level within high-dimensional contingency tables. This methodology represents an extension of our previous techniques and employs a linear index approach. The method simplifies complex, high-dimensional genetic data into a one-dimensional format, thereby improving computational efficiency while preserving the intricate structure of gene interactions. Theorems are developed which demonstrate how the sum of micromutual information asymptotically converges to the total MI for multidimensional variables. Our findings indicate that the maximum value of micromutual information, termed MIfront, adheres to an extreme value distribution. The observation of MIfront provides a streamlined approach to estimating the total MI, due to the simplicity of measuring the micromutual information of just one cell. This approach has the potential to improve data analysis in genomics and other fields that deal with multidimensional information.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.