Inconsistency, Instability, and Generalization Gap of Deep Neural Network Training

As deep neural networks are highly expressive, it is important to find solutions with small generalization gap (the difference between the performance on the training data and unseen data). Focusing on the stochastic nature of training, we first present a theoretical analysis in which the bound of generalization gap depends on what we call inconsistency and instability of model outputs, which can be estimated on unlabeled data. Our empirical study based on this analysis shows that instability and inconsistency are strongly predictive of generalization gap in various settings. In particular, our finding indicates that inconsistency is a more reliable indicator of generalization gap than the sharpness of the loss landscape. Furthermore, we show that algorithmic reduction of inconsistency leads to superior performance. The results also provide a theoretical basis for existing methods such as co-distillation and ensemble.

Paper

References (42)

Scroll for more · 30 remaining

Similar papers

Peer review

Reviewer TByp5/10 · confidence 4/52023-07-04

Summary

This manuscript propose new notions of inconsistency, instability, and information-theoretifc instability based on the output confidence score to estimate the generalization gap of deep neural networks. Theoretical and empirical results are presented and show that the proposed notions, especially inconsistency, are correlated well with well-trained neural networks.

Strengths

1. Three novel notions are proposed to measure the generalization gap of neural networks; 2. Extensive experiments have been conducted to verfify the good correlation of instability and inconsistency with the generalization gap of neural networks; 3. The manuscript is well organized and the writing is clear; 4. Emprical study show the better correlation of the inconsistency than disagreement.

Weaknesses

1. The novelty of proposed measurements are limited: - The Inconsistency and Instability are very similar to the definition of disagreement, while the former notions replace the outout from one-hot predictions to softmax confidence score. - The Instabilty of model parameter distributions ($\mathcal{I}_P$) can be regarded as a kind of algorithm stability. Therefore, I suggest the authors to provide a discussion on the differentce or advantages w.r.t. former definition of algorithm stabiliy. 2. Marginal contributions on the theoretical results (Theorem 2.1): - The upper bound in Theorem 2.1 is not a uniform convergence-based generalization bound, and does not show the relation to training sample size and may be loose, which greatly undermines the significance of the theoretical results; - The right hand of the given upper bound is hard to estimate due to the existence of the Instabiliyt of model parameter distributions $\mathcal{I}_P$, although In consistency $\mathcal{C}_P$ and Instability $\mathcal{S}_P$ can be convenient to estimate on unlabeled data. Based on above weaknesses and considering the marginal contributions on the proposed notions and theoretical results, I think this work is slightly below the acceptance bar.

Questions

See weaknesses.

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

See weaknesses.

Reviewer 3GRr6/10 · confidence 4/52023-07-05

Summary

The paper presents two measures for a stochastic training algorithm: inconsistency and instability. The former measures the inconsistency (or "disagreement") within the random ensemble of models trained from the same training set. The latter measures the inconsistency of two ensembled predictors, each obtained from an ensemble trained from an independent training set. A theorem is presented stating that the sum of inconsistency and instability modulates the mutual information (between the training set and algorithm output) in the generalization bound of Xu/Raginsky' 2017. Empirical investigation is performed to assess the predictiveness of inconsistency and instability for generalization gap. Algorithmic implications are also investigated.

Strengths

To this reviewer, a particularly novel and interesting aspect of this work is bringing the notion of inconsistency, or "within-training-set agreement" into the landscape of generalization bounds. This notion is akin to the notion of "generalization disagreement equality" (GDE) in the work of Jiang et al 2022 (reference [17] of this paper). Notably -- although not adequately discussed by the authors -- a sufficient condition of GDE is a notion of calibration in [17]. A potential impact of this work is extending the development of information-theoretic generalization bounds to include the calibration-alike quantities. This, I found intriguing and inspiring. The theoretical development is light. Nonetheless interesting. Among various empirical results, the most interesting and novel aspect to this reviewer is the observation that inconsistency is more predictive for generalization than sharpness. Algorithmic exploitation of this aspect is also interesting.

Weaknesses

Some conclusions from the empirical study appear speculative to this reviewer. For example, the authors hypothesized that a low degree of randomness is required for inconsistency+instability to be predictive, and in other places the author attributed complex phenomena arising from experiments to the interaction with the mutual information term. It is desirable that such claims are better corroborated.

Questions

1. Are the notions of inconsistency and instability related to the notion of functional-CMI (or functional MI) in the work of Harutyunyan et al, "Information-theoretic generalization bounds for black-box learning algorithms", NeurIPS 2021? Exploring this connection might enhance this work. 2. What is the loss function used in evaluating generalization gaps in the experiments?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

2 fair

Limitations

Nothing to add.

Reviewer k85B4/10 · confidence 4/52023-07-06

Summary

This paper investigates the generalization gap in deep neural networks, and propose that this gap is influenced by the inconsistency and instability of model outputs, two quantities which are defined by the authors, and justified theoretically via a new information theoretic generalization bound. The authors conduct empirical studies that confirm the predictive power of inconsistency and instability on the generalization gap, and demonstrate that explicitly reducing inconsistency during training improves performance.

Strengths

The paper is generally well-written, and includes several interesting observations. While I'm unsure of the novelty of Theorem 2.1, its form is compelling and I like the fact that $\mathcal{D}_P$ and $\mathcal{C}_P$ can be estimated efficiently (unlike the mutual information term $\mathcal{I}_P$ that also appears in other information theoretic bounds). I also find the results at the end on explicitly encouraging consistency interesting and a strong contribution -- in my opinion it would be good to expand on this aspect of the paper.

Weaknesses

While the form of the bound in Theorem 2.1 is compelling, I have some concerns with its novelty/improvement relative to existing results (see questions below). In particular, it's unclear to me that the bound represents an improvement on existing information theoretic generalization bounds in the literature. On the empirical side, I think a more rigorous correlation analysis of the $\mathcal{C}_P/\mathcal{D}_P$ metrics along the lines of prior work (e.g. Jiang et al. 2020) would help strengthen the claims of the paper, since the empirical correlation between these measures and the generalization gap seems to be one of its main contributions. I also think it would be helpful to have a more detailed comparison of the inconsistency/error relationship compares with the disagreement/error relationship observed in prior work, as these seem very closely related. Overall, while the paper is well-written and contains some interesting insights, at the current stage I think it lacks sufficient novelty/improvement on the theoretical side, and comparison to prior work on the empirical side to recommend acceptance. **References** Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, Samy Bengio, Fantastic Generalization Measures and Where to Find Them, 2020.

Questions

- How does the bound Theorem 2.1 compare with the large existing literature of information-theoretic generalization bounds? Perhaps the simplest of these from Xu & Raginsky (2017) is of the form $\sqrt{2\sigma^2 I/n}$, upon which your result seems like a small improvement at best, and only when $\mathcal{D}_P$ is very small. More recent bounds (e.g. Steinke and Zakynthinou, 2020 using the conditional mutual information) have made significant improvements on this, and so clarification as to which regimes your result provides an improvement would be helpful. - Relatedly, is there any evidence that $\mathcal{D}_p$ is the dominant term in Theorem 2.1? Previous work has noted that mutual-information based bounds can be extremely large (though they are difficult to numerically estimate), and seemingly here the term $\mathcal{I}_P$ would dominate. - Could the authors clarify what is being varied in the plots in, e.g. Figures 1 and 2? Are these plots illustrating the generalization gap/$\mathcal{D}_P$ in time, i.e. as a function of the iteration of optimization? If so, it would be useful if there was some indication of the direction of time in this plot. - Do the authors have any hypotheses for why the disagreement/error relationship exhibits different behavior than the generalization gap/inconsistency relationship (maybe specifically in the zero training error regime, in which the generalization gap = test error)? It seems like this may have to do with the distribution of the logits in the trained models (which is ignored by the disagreement but not by the inconsistency metric), which would be very interesting to understand. - Could the authors explicitly state the penalty that is added to the loss function during training to encourage consistency and explain how it is computed? **References** Thomas Steinke and Lydia Zakynthinou, Reasoning About Generalization via Conditional Mutual Information, 2020.

Rating

4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

4 excellent

Presentation

4 excellent

Contribution

3 good

Limitations

N/A

Reviewer ARdH6/10 · confidence 3/52023-07-07

Summary

In this work, the authors introduce the ideas of “instability” and “inconsistency” of model outputs, and investigate the relationship between these quantities and the generalization gap. In particular, they empirically find a positive correlation between instability + inconsistency and the generalization gap. They further show that inconsistency can be more predictive of the generalization gap than m-flatness.

Strengths

1. The paper is written cleary, well–organized, and well-motivated. The authors do a good job in providing some intuition behind the mathematical definitions of inconsistency and instability. 2. The paper includes extensive experiments to support their hypothesis (e.g., use of various architectures, vision and text datasets, training methods, etc.)

Weaknesses

1. In Figure 2, the authors observe that $\mathcal{D_P}$'s predictive ability of generalization gap is sensitive to the learning rate. The authors further suggest that when the “final randomness is high,” $\mathcal{D_P}$’s predictive ability of the generalization gap is not as strong. In practice, practitioners may use a large learning rate and small batch size to obtain well-generalizing models (and so the final randomness would be high in this situation). Thus, $\mathcal{D_P}$ may not be useful here. (In some sense, it is not surprising that low final randomness correlates with $\mathcal{D}_{P}$’s predictive ability of generalization gap). 2. The authors choose a constant learning rate schedule for the starting experiments to avoid confounding variables. However, the architectures used in the experiments have normalization layers, which can induce learning rate schedules. Perhaps running the preliminary experiments with at least one architecture w/out any normalization layers would be beneficial. 3. The benefit of $\mathcal{D}_{P}$ over prior metrics such as disagreement is not evident. In [17], test error is calculated at the end of training (when the train error is nearly zero). Since this is not true in figure 10, I would be interested in a plot of the gap between train accuracy and test accuracy vs. training loss instead of figure 10d. (aside: the gap between train accuracy and test accuracy is an alternate definition of generalization gap to the definition used in this work). Perhaps this would lead to a more appropriate comparison of inconsistency vs. disagreement detailed in lines 113-127. More generally, it would be nice to also include plots for this alternative definition of generalization gap, especially since experiment performance in the paper is often measured via test error. 4. In the appendix (lines 600-604), $\rho$ was set based on reference or the development data. However, the optimal value of $\rho$ can be quite sensitive to the architecture used. Thus, it would be nice to use some type of grid search to choose $\rho$ for fair comparison in e.g., figure 8.

Questions

It would be helpful to include legends in each figure (e.g., figure 4 is missing a legend. I assume the legend is consistent with figure 2. However, it would be convenient to include the legend again.) What are the training errors for models in each experiment?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

The authors have described the limitations.

Reviewer ARdH2023-08-21

I acknowledge and appreciate the authors' responses. I intend to keep my score.

Reviewer 3GRr2023-08-22

Thank you for the reply.

I will keep the rating.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC