Beyond Normal: On the Evaluation of Mutual Information Estimators

Mutual information is a general statistical dependency measure which has found applications in representation learning, causality, domain generalization and computational biology. However, mutual information estimators are typically evaluated on simple families of probability distributions, namely multivariate normal distribution and selected distributions with one-dimensional random variables. In this paper, we show how to construct a diverse family of distributions with known ground-truth mutual information and propose a language-independent benchmarking platform for mutual information estimators. We discuss the general applicability and limitations of classical and neural estimators in settings involving high dimensions, sparse interactions, long-tailed distributions, and high mutual information. Finally, we provide guidelines for practitioners on how to select appropriate estimator adapted to the difficulty of problem considered and issues one needs to consider when applying an estimator to a new data set.

Paper

Similar papers

Peer review

Reviewer 62uG6/10 · confidence 4/52023-06-22

Summary

The authors propose a method for creating expressive distributions via injective mappings that maintain their original MI. They state this in Theorem 2.1 and prove it in the Appendix. In addition to this, the authors benchmark a variety of estimators for MI on a set of tasks, including high-dimensional, long-tailed distributions, sparse interactions, and high mutual information. This information is contained in Sections 3-4, where they describe each task and critique each estimator for the tasks. From these experiments, the authors provide guidelines for practitioners on how to choose an estimator in Section 6.

Strengths

The paper is clear, provides many novel benchmarking tasks, and is easily verifiable. The results are reproducible as the authors have shared their code and documented their experimental parameters.

Weaknesses

1. The author does a poor job of motivating each data setting and explaining why each one is important. It would be beneficial to provide examples of domains where long-tail distributions are common, such as insurance, and elaborate on the significance of the other data settings as well. 2. I believe the last row of Figure 2, "True MI," could be better highlighted as it was difficult to discern that it was the point of comparison. 3. The main contributions of the paper seem relatively minor, as the authors are primarily considering more data settings than previous work when comparing MI estimators. However, addressing the first point could help alleviate this concern. 4. The main theoretical result, Theorem 2.1, appears to have already been demonstrated in the appendix of the following paper: https://arxiv.org/pdf/cond-mat/0305641.pdf. 5. The authors seem to have switched \citep for \cite in their paper.

Questions

Why not utilize the result of Theorem 2.1 and apply it to Normalizing Flows? I believe this combination would enhance the paper's distinctiveness. How is Theorem 2.1 different from that of the result in the appendix of https://arxiv.org/pdf/cond-mat/0305641.pdf?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

The authors have addressed the limitations of their work.

Reviewer vCqe7/10 · confidence 3/52023-06-23

Summary

This paper identifies a clear problem with many mutual information estimation benchmarks: most of them focus on simple (normal) distributions. The authors present a new set of forty (!) tasks that contain ground-truth informations, which can be constructed by noting that only injectivity is needed for an information-preserving transformation. The authors identify four distinct challenges: interaction sparsity, long tails, invariance, and high information values. Several conclusions about existing estimation algorithms are then made regarding the extensive analysis.

Strengths

- I enjoyed reading this paper. It clearly defines a goal and identifies key problems with existing approaches. - The paper presents mathematical background in a precise and effective way. - The structure of the paper is clear, using figures for clarifications where needed. - The paper is well-written with clear sentence structures and no grammatical or spelling errors. - The analysis is thorough, presenting forty distinct MI estimation tasks. - I enjoyed how instead of simply presenting the results, the authors dug deeper and identified four distinct challenges for MI estimators.

Weaknesses

- One could say that the paper lacks slightly in terms of originality and contributions.

Questions

- Out of curiosity: could modern invariant neural network architectures be used to obtain MI estimates invariant to diffeomorphisms? ### Conclusions While the overall contribution could be limited in terms of a model development sense, I think the paper identifies serious issues with modern MI estimation benchmarks. The paper not only provides new benchmarks that address these issues but also makes an effort to identify what aspects of MI estimation can make the problem hard. I foresee much new research originating from the identification of these aspects, where future papers focus in on them and propose methodologies that overcome these challenges. On top of that, the paper is very well written. Hence, I would recommend acceptance.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

4 excellent

Presentation

4 excellent

Contribution

3 good

Limitations

The authors clearly discuss the limitations of their study.

Reviewer FjRx5/10 · confidence 4/52023-06-30

Summary

A test benchmark for the evaluation of mutual information estimators is established and many different estimators compared. The test cases contain student-t and normal distributionas and their injective transformations. Difficult cases are discussed and evaluated in more detail.

Strengths

The code is reproducible and thus might be used for other estimators in the future The paper is well-written and easy to understand. It is important to make the community aware that evaluation of MI estimators on Gaussian distributions is pointless as these only depend on the covariance structure, so the big plus of the MI that it goes beyond te linear dependencies is ignored. The paper makes a strong point here by including a simple covariance estimator as well. The results on heavy tails are particularly interesting.

Weaknesses

The choice of the distribution used is not sufficiently argued. In particular, it is known that no MI estimator can evaluate MI correctly on arbitrary distributions. Only with restriction to a class of probability distributions (e.g., probability density functions with Lipschitz constraints) there is hope that estimation works. It is thus quite pointless to evaluate MI estimators on what seems like an "educated guess" of diverse distributions. The authors seem to not be aware of the huge theoretic background of MI, e.g., that arbitrary measureable injective mappings do not change MI and thus their only Theorem 2.1 is well-known in a much more general setting. This can be argued by the data processing inequality in two directions (X-Y-f(Y) and X-f(Y)-Y are both Markov chains) or directly the definition of MI via countable partitions that do not change if we use a measureable injective mapping.

Questions

What was the reason for choosing this specific set of distributions? What is "pointwise MI"?

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

4 excellent

Presentation

4 excellent

Contribution

2 fair

Limitations

The authors state that there are more interesting cases and they only cover transforms of normal and student-t distributions. They also mention that prior information might be incorporated. However, as said above it is known that an MI estimator can be fooled arbitrarily (giving any value for the MI) if one can choose the distribution freely. Thus, prior information is also included in the test cases here it is just not mentioned explicitly.

Reviewer AkdF6/10 · confidence 3/52023-07-07

Summary

This paper focuses on the topic of mutual information and shows how to construct a diverse family of distributions with known ground-truth mutual information. It's worth noting that obtaining a closed-form solution for mutual information is highly dependent on the specific assumptions and functional forms used for the variables X and Y. In practice, deriving closed-form expressions for mutual information can be challenging and may require additional simplifying assumptions or specific knowledge about the distributions involved. In contrast to previous works that typically assess mutual information estimators using simple probability distributions, this paper introduces a novel approach to constructing a diverse family of distributions with known ground-truth mutual information. Additionally, the authors propose a language-independent benchmarking platform to assess mutual information estimators. The authors explore the applicability of classical and neural estimators in scenarios involving high dimensions, sparse interactions, long-tailed distributions, and high mutual information. By examining these challenging settings, they provide insights into the strengths and limitations of different estimators. Moreover, the paper offers guidelines for practitioners to select the most suitable estimator based on the specific problem's difficulty and considerations when applying an estimator to new datasets. By presenting a comprehensive evaluation framework and practical recommendations, this research aims to advance the understanding and application of mutual information estimation in various domains.

Strengths

The mutual information estimator is an essential tool in causality, it can help discover the underlying causal graph or inference the strength of causal relations. However, as I mentioned earlier, deriving closed-form expressions for mutual information can be challenging in practice and may require additional simplifying assumptions or specific knowledge about the distributions involved. The paper introduces a method to construct a diverse family of distributions with known ground-truth mutual information. This is a significant contribution as it allows researchers to explore and evaluate mutual information estimators across various scenarios, encompassing various data characteristics and relationships. For example, explore gene regularity networks, understand the treatment effect in medical health care, gain insight for constructing a recommendation system, etc. The research paper presents a comprehensive evaluation framework for mutual information estimation, encompassing the construction of diverse distributions, benchmarking platform, exploration of challenging scenarios, and practical guidelines. This framework provides a holistic view of the estimation process, aiding researchers and practitioners in understanding, comparing, and selecting mutual information estimators effectively. Furthermore, the authors investigate the applicability of classical and neural estimators in challenging scenarios involving high dimensions, sparse interactions, long-tailed distributions, and high mutual information. This exploration provides valuable insights into the performance, strengths, and limitations of different estimators under these challenging conditions, enhancing our understanding of their effectiveness in real-world settings.

Weaknesses

1. Not all joint distributions can be represented in the form of $P_{f(x)g(x)}$, limiting the applicability of the benchmark to a specific set of distributions. Extending the family of distributions with known mutual information and efficient sampling is seen as a natural direction for future improvement. 2. Even though the benchmark demonstrates that distributions with longer tails pose a harder challenge for the considered estimators, applying a transformation like the asinh transform does not fully address the issues. 3. The summary does not mention external validation or comparisons between the proposed approach or estimators and existing methods or benchmarks in the field. The absence of such external validation makes it difficult to assess the generalizability or superiority of the contributions in relation to established techniques or alternative approaches.

Questions

See above "weaknesses".

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

See above "weaknesses".

Reviewer FjRx2023-08-11

Thank you for the detailed response. Indeed, I had a hard time finding a source for Theorem 2.1 as well and can only provide data processing inequalities that then imply the authors result as corollary. I also have to admit that mere measurability as conjectured in my original review is not sufficient and some way to show measureability of the inverse (defined on the range) is also required. I'm still reluctant to support the given set of distributions as reasonably representative but have to admit that so far MI has been estimated on much worse datasets, thus it is a step in the right direction. I also question the statement that "an analytical expression for non-zero ground-truth MI is currently tractable only for the multivariate normal and Student families." With some basic math, other approches like sums of uniform distributions should also be tractable. Nevertheless, I'll increase my score as the paper is a step in the right direction but hope that soon more versatile distributions will be added to the benchmark and we are not stuck with a less than perfect solution for the next decade.

Authorsrebuttal2023-08-14

Thank you very much for your response. We agree that the chosen set of distributions is not universal, but, at the same time, is a step in the right direction. Regarding the sums of uniform distributions, we would like to thank you for the suggestion. We have already included in our benchmark the bivariate case (lines 134–137): $Y = X+N$, where $X\sim \mathrm{Uniform}(0, 1)$ and $N\sim \mathrm{Uniform}(-\varepsilon, \varepsilon)$, but we agree that a multivariate generalization (with independent $X_1, \dotsc, X_k$ and $N_1, \dotsc, N_k$) has tractable ground-truth mutual information as well. We will add it to the benchmark.

Reviewer vCqe2023-08-16

Thank you for providing the detailed response. Your elaboration on the potential applications and limitations of invariant neural networks in relation to diffeomorphisms is enlightening. Thanks for the references.

Reviewer 62uG2023-08-16

Thank you for your response. Theorem 2.1: I appreciate your effort in providing further clarification. Nevertheless, I maintain my belief that the novelty of the theorem might still appear ambiguous to readers. To enhance its clarity, I suggest considering referencing the proof by Kraskov et al. (2004) or another similar result. Despite its familiarity within the community, it's important to remember that individuals from outside the community might perceive this as novel. Motivating data settings: I commend the authors for their thorough approach in encompassing a diverse range of distributions. Thus making their benchmark applicable to a large set of domains.. My primary reason for the initial low score pertains to the treatment of Theorem 2.1. If the authors could furnish additional context surrounding Theorem 2.1, I would happily raise my evaluation. I appreciate the authors for their considerate response and answers to all my questions.

Authorsrebuttal2023-08-17

Thank you for your suggestion and encouraging words on the diverse range of distributions covered! We have now understood your argument regarding Theorem 2.1 and we fully agree that adding more context to it will significantly increase manuscript clarity. We will add the following paragraph to Section 2: > Theorem 2.1 is a well-known property of mutual information, formulated in various versions. For example, Kraskov et al. (2004) consider a case in which $f$ and $g$ are diffeomorphisms and all measures have probability density functions. For the sake of completeness, we include a proof of Theorem 2.1 (covering singular measures and any continuous injections) in Appendix A. Please, let us know if you have any further suggestions.

Reviewer 62uG2023-08-17

Thank you for the change, I have changed my score to weak accept. Good luck!

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC