Towards the Dynamics of a DNN Learning Symbolic Interactions

This study proves the two-phase dynamics of a deep neural network (DNN) learning interactions. Despite the long disappointing view of the faithfulness of post-hoc explanation of a DNN, a series of theorems have been proven in recent years to show that for a given input sample, a small set of interactions between input variables can be considered as primitive inference patterns that faithfully represent a DNN's detailed inference logic on that sample. Particularly, Zhang et al. have observed that various DNNs all learn interactions of different complexities in two distinct phases, and this two-phase dynamics well explains how a DNN changes from under-fitting to over-fitting. Therefore, in this study, we mathematically prove the two-phase dynamics of interactions, providing a theoretical mechanism for how the generalization power of a DNN changes during the training process. Experiments show that our theory well predicts the real dynamics of interactions on different DNNs trained for various tasks.

Paper

References (46)

Scroll for more · 34 remaining

Similar papers

Peer review

Reviewer xnzG8/10 · confidence 3/52024-06-19

Summary

This paper studies the training dynamics (underfitting to overfitting) of deep neural networks via the perspective of symbolic interactions. They formulate the learning of interactions as a linear regression problem on a set of interaction triggering functions. They show the two-stage dynamics. In the first stage, neural networks first remove initial interactions and learn low-order interactions. In the second stage, neural networks learn increasingly more complicated interactions, leading to overfitting. Through empirical experiments, they show the story is valid for various architectures.

Strengths

* The paper is well written, clear in motivation. Figures are pleasure to read. * Statements are justified with math theorems. * The link between underfitting-overfitting vs the distribution of I_k is novel. * Testing extensively on various architectures

Weaknesses

* The contribution is not fully clear. Since this work heavily relies on [26][27][45], it should clearly highlight what's new and what's known in pervious works.

Questions

* Line 21, "the three conditions". what are these three conditions? * Line 21, "for" -> "For" * Line 100, Eq. (2), can you explain more where does the factor (-1)^{|S|-|T|} comes from? This sign seem to be canceling terms out? * In Figure 2, it would be nice to also include training loss, not just train-test gap. Is there a time point when the interactions are too simple such that increasing order actually helps alleviate underfitting but still not overfitting? * How does the initial I_k distribution change with different initialization scales?

Rating

8

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

The authors adequately addressed the limitations.

Reviewer bep27/10 · confidence 3/52024-07-12

Summary

This study investigates the two-phase dynamics of DNNs learning interactions during training, demonstrating that DNNs initially focus on simpler, low-order interactions and progressively transition to more complex, high-order interactions. The learning process is reformulated as a linear regression problem, which facilitates the analysis of how DNNs manage interaction learning under parameter noise.

Strengths

- The research backs its theoretical claims with sufficient experimental evidence. This thorough analysis strengthens the credibility of the study's conclusions. - The study explores the two-phase dynamics of DNNs, clarifying how networks transition from simple to complex interactions. This clarifies the mechanisms underlying neural network generalization and susceptibility to overfitting.

Weaknesses

- The experiments on textual is limited compared to those on vision tasks.

Questions

- How is it ensured that the model does not re-learn the initial interactions during the second phase? Please justify. - Please justify the disparity in VGG-16 on CIFAR-10 and VGG-11 MNIST based on Figure 5. It seems their distribution is a bit different from the others in the last time point.

Rating

7

Confidence

3

Soundness

3

Presentation

3

Contribution

2

Limitations

- The experiments on textual data are limited. Expanding the analysis to include a broader range of datasets could provide a more comprehensive understanding of the DNN's interactions.

Reviewer iitK5/10 · confidence 2/52024-07-13

Summary

The paper investigates the two-phase dynamics of interactions during training by reformulating the learning of interactions as a linear regression problem. The authors provide an analytic solution to the minimization problem and use this solution to explain the two-phase dynamics of interactions.

Strengths

The paper provides **a theoretical explanation for the two-phase dynamics of interactions**, which appears to be a novel contribution to the literature.

Weaknesses

* **Presentation of the paper is difficult to follow**, especially for someone like me who is not familliar with the literature. It would benefit from improved clarity and readability, particularly in summarizing relevant literature and explaining the mathematical setting. * The paper only characterizes minimizers, which are the ending points of learning and **do not consider the exact training dynamics**.

Questions

* Do you think **it is possible to analyze the exact dynamics** of interactions by considering specific neural network architectures and data models? This direction may enhance the results of the paper. * Are there any **practical implications** of the theoretical findings?

Rating

5

Confidence

2

Soundness

2

Presentation

1

Contribution

2

Limitations

N/A

Reviewer bep22024-08-09

The authors have clarified my questions. I would recommend that the authors include the explanation given for "Q2: re-learning initial interactions" in the final version of the paper if accepted. I have decided to raise my score.

Authorsrebuttal2024-08-10

Thank you very much. We will follow your suggestion to incorporate the explanation for "Q2: re-learning initial interactions" into the paper if the paper is accepted.

Reviewer iitK2024-08-11

Thank you for your detailed response. The response adequately addressed my concerns. After reviewing the discussion between the authors and reviewers and re-reading the draft, I gained a better understanding of the work. As a result, I have decided to increase my score. I hope the authors will further enhance the paper by incorporating the more detailed background and discussion points mentioned in their response. I believe these improvements will make the paper more accessible to readers who are less familiar with this literature.

Authorsrebuttal2024-08-12

Thank you for very much. We will follow your suggestion to provide a more thorough review of the background and incorporate the discussion points in the next version of the paper if accepted.

Reviewer xnzG2024-08-11

I want to thank the author for addressing my concerns. I'm more convinced now this is an important paper for its novel concepts. I'll raise my score to 8.

Authorsrebuttal2024-08-12

Thank you very much for your appreciation. We will continue to enhance the paper according to the discussion with all reviewers. We hope this paper provides deep insights into the two-phase dynamics of interactions and its tight connection to the generalization power of DNNs.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC