The representations of neural networks are often compared to those of biological systems by performing regression between the neural network responses and those measured from biological systems. Many different state-of-the-art deep neural networks yield similar neural predictions, but it remains unclear how to differentiate among models that perform equally well at predicting neural responses. To gain insight into this, we use a recent theoretical framework that relates the generalization error from regression to the spectral properties of the model and the target. We apply this theory to the case of regression between model activations and neural responses and decompose the neural prediction error in terms of the model eigenspectra, alignment of model eigenvectors and neural responses, and the training set size. Using this decomposition, we introduce geometrical measures to interpret the neural prediction error. We test a large number of deep neural networks that predict visual cortical activity and show that there are multiple types of geometries that result in low neural prediction error as measured via regression. The work demonstrates that carefully decomposing representational metrics can provide interpretability of how models are capturing neural activity and points the way towards improved models of neural activity.
Paper
Similar papers
Peer review
Summary
This submission uses spectral theory and simulation experiments to try to compare and assess the representations of DNNs vs those of biological neural networks.
Strengths
The paper comes with a fairly extensive review of the literature. The paper uses both theoretical ideas and simulations.
Weaknesses
The paper is difficult to read and understand. The primary reason is that the problem being tackled is not stated in any clear way. The boundary between artificial NNs and biological NNs is very fuzzy throughout the paper and it is often unclear if the authors are discussing biological NNs or artificial NNs. The results also are not very clear. This is already obvious from the abstract which does not contain a single quantitative result. Furthermore, there are too many concepts piled up in the paper in a messy way (spectral, geometry, alignment, adversarial vs standard training, error mode weights, etc.). The authors cite 43 papers. If the goal is to provide a fairly complete overview of the relevant work, then it seems that some key papers are missing (e.g. Anderson and Zipser in the 1980s, more recent work by DiCarlo et al. )
Questions
There may be good ideas behind this paper and in time this direction of research may prove useful. But at this stage things look premature and very foggy. In future versions, it will be essential to improve the clarity and keep very clear distinctions between artificial and biological NNs. A term like "neural predictivity" is extremely ambiguous in your context as it could have several meanings.
Rating
4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Soundness
2 fair
Presentation
2 fair
Contribution
2 fair
Limitations
In addition to the problems mentioned above, there are significant limitations posed by the size of the data sets. The authors are aware of this problem and mention it in Section 5 together with other limitations.
Summary
The paper uses (but doesn't introduce) a theoretical framework that relates generalization error to spectral bias of network activations to geometrically/spectrally analyze what aspects of pretrained deep networks contribute to the predictive performance of neural activity in layers V1, V2, V4 and IT. The authors verify that the theory matches empirical values, and analyze different aspects of the generalization, such as the influence of the training data size and adversarial pretraining.
Strengths
It is inherently difficult to analyze what aspects of deep network representations contribute to the predictive performance on neural data. The authors make an important contribution that disentangle different factors. This will help to understand why and when deep representations predict neural data well.
Weaknesses
My main two criticisms are: - The clarity/accessibility could be improved at times. - There are a few more references that could be considered. - Some of the assumptions/limitations could be discussed more clearly. While I like the general approach, I think it could be clearer at some point. Since you, as the authors, have thought a long time about what the different $W_i$, $R$, $D$ and other values mean, the reader sees that for the first time. So I think it's important to give a good intuition about what they mean. Figure 1 didn't really help in that respect. Particularly, D and E need a bit more effort to be helpful. The reason why this is so crucial is that they mainly discuss the observations in terms of these values. So if the reader doesn't have a clear intuition what they mean, the observations/interpretations become meaningless. So I urge the authors to improve on the introduction part and on the observation part (--> Q1). I think the following papers would be worth considering to cite (no, I am not S. Cadena ;) ): * Santiago A. Cadena, Konstantin F. Willeke, Kelli Restivo, George Denfield, Fabian H. Sinz, Matthias Bethge, Andreas S. Tolias, Alexander S. Ecker Diverse task-driven modeling of macaque V4 reveals functional specialization towards semantic tasks * Santiago A. Cadena, Fabian H. Sinz, Taliah Muhammad, Emmanouil Froudarakis, Erick Cobos, Edgar Y. Walker, Jake Reimer, Matthias Bethge, Andreas Tolias, Alexander S. Ecker How well do deep neural networks trained on object recognition characterize the mouse visual system? * S. A. Cadena, G. H. Denfield, E. Y. Walker, L. A. Gatys, A. S. Tolias, M. Bethge, and A. S. Ecker Deep convolutional models improve predictions of macaque V1 responses to natural images The first, because it discusses different prediction performances depending on the pretrained representation. The second because it finds that random representations do almost equally well as pretrained (albeit for mouse). And the third, because it compared task driven and data driven prediction performance for monkey V1. As far as I know, the authors also provide the data. So you could even repeat some of the analysis with more data for V1 (135 texture stimuli are not a lot). Re Limitations/Assumptions: Some of the assumptions seem unrealistic to me. I would ask the authors to discuss that in more detail. Specifically - Q2: You write that model and brain responses need to be deterministic functions of the stimuli. For brain responses this is certainly not true since there is noise and also latent brain states. Can you discuss how the deviation from this assumption affects your results. - Q3: I am unsure about the $M,P\rightarrow \infty$ and $M/P \in O(1)$ assumption. This means that the model features need to go to infinity as the data grows. Doesn't that mean that the model has to be non-parametric, which is not true for deep network. I would ask the authors to discuss that in more detail. Discussing the limitations in two very short paragraphs in section 5 is not enough. Especially, since they have two pages left.
Questions
Q1-Q3: See above. Q4: I couldn't find how you extracted the 516 model activations (neither in the paper nor in the supp material). Please briefly describe this procedure in the main paper (i.e. did you use PCA? on how many activations/stimuli? did you only use one layer or many?). Minor: - l 97: "self-consistently" did you mean "self-consistency"?
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Soundness
3 good
Presentation
2 fair
Contribution
3 good
Limitations
See weakness above.
Summary
This study asks a crucial question: What is it about representations in ImageNet-trained neural networks that allow the successful regression of responses in mammalian visual cortex? The authors answer this question (in part) via an ingenious method combining learning theory and empirical analyses. Even after a decade of study (or more), it is unclear why the responses of neurons in the visual cortex can be so well predicted – better than any other approach – by linearly regressing them from the representations of pretrained DNNs. Here, the authors open the door to a potentially very fruitful method of answering why. Beginning from an analytical expression for the generalization error of ridge regression, the authors notice that it can be decomposed into a product of two terms describing the geometry of the error in the space of examples (the outer product a.k.a. Gram matrix space): one term which carries the interpretation of an average radius, and the other which carries the interpretation of the effective dimensionality of the error. A low generalization error can be achieved by decreasing this radius and dimension. Then, the authors empirically (and thoroughly) analyze how the generalization error decomposes into radius and dimension when predicting V1, V2, V4, and IT responses from a very wide range of pretrained networks, both supervised and unsupervised. The results are striking. For example, networks appear to be surprisingly constrained in their “error mode geometry”, apparently forced to trade between low radius and high dimension. There are many implications in this paper for how we go about understanding neural responses.
Strengths
A pleasure to review. This paper tackles a question of high significance using a thorough and deep analysis, one that is both highly original and leverages recent advances. The empirical analyses are complete and leave nearly nothing wanting.
Weaknesses
The only weakness I might comment on, in the category of clarity, is the absolute density of results, many of which are not followed up or explored more deeply. To help the reader internalize these interesting results, it would be better to provide a writing structure with more summaries of the outcomes of interest both at the paragraph level and at the paper level. (I personally prefer the CCC style argued for by Mensh and Kording (2017)). At the moment the reader is left to determine for oneself which results to find interesting – meaning that to many quick readers some of these important lessons will be overlooked. I have a number of requests for expanding the exposition. Apologies, as not all of these will fit in the page limit! First, it would be nice to have a little more discussion and intuition of the meaning of the key error mode geometry terms, radius and dimension. Then, there were a number of interesting findings that were only given a sentence or not mentioned in the main text. - SI.5.4 is very interesting but not much mentioned in the main text. Consider moving to the main text; by way of contrast this might help distinguish your own definition of dimension. - Line 443 in the supplement, which mentions that the eigenspectra of the stimulus set in part controls generalization error, which means that it could be a tool for better selecting stimulus sets for neural predictions. What would this look like? - It would be good to mention in the main text (briefly) that you analyze both supervised and unsupervised networks, and of many architectures. - That “the improved predictivity of trained networks in regions V2, V4, and IT is primarily driven by changes in their intrinsic expressivity as summarized by their eigenspectrum decay, rather than significant changes in their alignment with neural data.” (160-162). This would have large implications for those believing the brain==ImageNet DNNs. Could this be unpacked? These are just a few; it would be nice to have summaries of all main takeaways.
Questions
Could a plot be created in which the error mode geometry is shown split by which objective function (supervised, Barlow Twins, etc) as well as each network? Even a null result (indistinguishable) would be interesting. Does the theory require an alignment between the eigenmodes of $G_R$ and $G_X$? If so, how is this achieved in theory?
Rating
9: Very Strong Accept: Technically flawless paper with groundbreaking impact on at least one area of AI/ML and excellent impact on multiple areas of AI/ML, with flawless evaluation, resources, and reproducibility, and no unaddressed ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Soundness
4 excellent
Presentation
4 excellent
Contribution
4 excellent
Limitations
The limitations were thoroughly mentioned and addressed. I greatly appreciated the exploration (in the supplement) of the difference between the generalization error of ridge regression and the Pearson’s R2 measure for partial least squares regression commonly used elsewhere. The additional supplement plots for other small details (outliers, train/test split, $l_\infty$ adversarial) were also appreciated. In general, I appreciate how the analyses and plots were chosen not to puff up results or push them into any particular narrative, but rather were an honest and complete presentation of the results.
Summary
Previous works have demonstrated that many different state-of-the-art deep neural networks (DNNs) perform similarly at neural responses prediction. But a complete understanding of which aspects of these DNNs lead to the similarity in predicting neural responses remains unknown. The authors proposes a spectral theoretical framework to explore how the geometrical properties of DNNs affect the performance of neural responses prediction. More specifically, the authors introduce a generalization error to serve as a measure of fit, which brings a geometrical interpretation of the model activations (representations) in DNNs. This geometrical interpretation could further give additional insights into how DNNs are achieving different performance of neural responses prediction. The authors also design several experiments to investigate the roles of layer depth, dataset size, and different training approaches of DNNs in neural responses prediction.
Strengths
* The authors establish a novel link betwwen the predictivity of neural responses and the geometry of DNN representations. * The authors provide a solid analysis and several comprehensive experiments to support their theoretical framework.
Weaknesses
* The authors should improve the clarity of their conclusions from their results. Section 3 provides many details about their analysis and results, but it's hard for the readers to extract the essential arguments the authors want to convey. It would be better to summarize your conclusions at the beginning or end of each paragraph. * The authors should improve their exposition of the spectral theoretical framework. I recommend the authors include more background information and provide some examples to show the role of each equation in the relationship between the generalization error and the geometry of model activations.
Questions
N/A
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Soundness
3 good
Presentation
2 fair
Contribution
3 good
Limitations
I have highlighted weaknesses above. I have nothing further to add here.
Summary
The authors studied linear regression-based DNN encoding models of macaque visual cortical areas by extending a recent generalization error theory. They provided a theoretical link between the predictivity and geometry of representations and showed that models with similar generalization errors may have quite different geometrical properties. They compared a randomly initialized ResNet50 with standard and adversarially trained ResNet50 models. Evaluations and analyses are a bit limited.
Strengths
- Analyzing the eigenspectrum of a model's feature space appears to be an interesting analysis that goes beyond the standard measures of representational similarity analysis or linear regression experiments that one usually comes across for this line of research at the intersection of cognitive neuroscience and machine learning. So, I appreciate that effort.
Weaknesses
**Major** - **Tense**: There's a mix between present tense and past perfect in $\S2$ which I think is not great. Choose one tense and use it consistently. I'd use the present tense to describe the problem settings and methods and either use present tense or past perfect for the results section. Also, use active rather than passive voice which I think is better readable (and more scientific). - **$\S3.4$ + main results**: I find it very strange to denote the training set size with small $p$. As I mention below, $p$ is usually used to refer to the number of features in some representation space and not to the size of the (training) data. If I was just reading the title of the subsection, I'd think that the generalization error depends on the size of the feature space (although I would not know on which of the feature spaces). This is not implausible, but very different from what you mean/find. Actually, it's pretty trivial that the generalization error depends on the size of the training data for fitting the regression model. When would the generalization error not depend on the size of the training data? This section is in general a bit confusing. In lines 205-207 you write: "[...] it may be that trained networks do better than their random counterparts at predicting V1 data for large, but not small, training data set sizes." Do you refer to the size of the (training) data with which the regression model is fit? If that's the case, I am a bit puzzled because from Bayesian statistics (and statistics in general) we know that the prior is most important for small data regimes. In the infinite data limit, the prior and, therefore, any regularization parameter stops mattering. A pretrained neural network has a much better initialization compared to a randomly initialized network. It has found a good local minimum over the course of pretraining which gives it a big head start. The pertaining is an implicit regularization for any downstream task. Here, we can treat the regression task as a downstream task. So, the prior should matter most when the training data to fit your regression model is small. You seem to find the opposite, or at least that's what you claim. The only possibility where I can see this being the case is when the training data size for the downstream task (e.g., for performing linear regression) is so small that the parameters of your (regression) model can basically explain no variance at all. But then the findings are pretty useless, aren't they? Actually, I am confident that $N=135$ is too small to find any statistically meaningful effects. - **Model choice**: It would have been interesting to see at least one, if not more, other architecture(s) being analyzed. I doubt that you can generalize any findings from a single ResNet50 to the whole space of DNNs. Even within the subclass of CNNs, there exist countless models. I find it more interesting to compare CNNs (pick two or three) to ViTs (pick two or three) but comparing a few CNNs (each with a different inductive bias --- e.g., AlexNet, VGG16, ResNet50, ResNext, ConvNeXt, EfficientNetV2), ideally trained with different recipes would have probably been sufficient. **Minor** - Lines 26-27: "[...] many different architecture and training procedures lead to similar predictions of neural responses." While the former is true and has been demonstrated multiple times in various ways, it has recently been shown in a comprehensive study that both training data and objective function have a major impact on the degree of alignment with human behavior; much more than architecture (see [Muttenthaler, L., Dippel, J., Linhardt, L., Vandermeulen, R. A., & Kornblith, S. (2022)](https://openreview.net/pdf?id=ReDQ1OUQR0X)). A longer while ago it has also been shown that different data augmentation strategies lead to different biases and as such to different degrees of consistency with human errors. See [Hermann, K., Chen, T., & Kornblith, S. (2020)](https://proceedings.neurips.cc/paper_files/paper/2020/file/db5f9f42a7157abe65bb145000b5871a-Paper.pdf); [Hermann, K., & Lampinen, A. (2020)](https://proceedings.neurips.cc/paper/2020/file/71e9c6620d381d60196ebe694840aaaa-Paper.pdf); [Geirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F. A., & Brendel, W. (2018)](https://arxiv.org/pdf/1811.12231.pdf); and [Geirhos, R., Narayanappa, K., Mitzkus, B., Thieringer, T., Bethge, M., Wichmann, F. A., & Brendel, W. (2021)](https://proceedings.neurips.cc/paper_files/paper/2021/file/c8877cff22082a16395a57e97232bb6f-Paper.pdf). - In line 58 is a small grammar mistake: "[...] characterizing how fast the eigenvalues of data Gram matrix fall [...]". There is an article missing before **the** data Gram matrix. - **Notation**: The notation in $\S2.2$ is weird. Why is the matrix of neural activations $\mathbf{R} \in \mathbb{R}^{P \times N}$? First, $\bf{R}$ can sometimes have a special meaning. So, I'd be careful with that. Second, although not bolded, you use $R$ for the generalization error terms. Third, generally $N$ is used to denote the number of stimuli/examples in your training data and not to refer to the number of features. My suggestion is to denote the matrix of model features as $\mathbf{X} \in \mathbb{R}^{N \times D}$ and the matrix of neural activations $\mathbf{Y} \in \mathbb{R}^{N \times P}$. As such, the regression problem $\mathbf{Y} = \mathbf{X}\mathbf{A}^{\top} + b$ becomes more intuitive. I'd replace $\mathbf{R}^{\ast}(\alpha_{reg}, p)$ with the full equation. $\mathbf{R}^{\ast}(\alpha_{reg}, p)$ is not easy to parse. That being said, note that this is just my personal taste which is why it's a minor weakness. There may be people who don't care about that. I think that maths should be easy to follow, even if a reader skims the paper. I dislike decorative maths. If it's just decorative, it can as well be omitted. - In Equation 2 is an "is equivalent to" symbol. It is probably not clear to everyone why the two terms are equivalent but not equal. I'd like to see a (short) derivation of it. You can derive it in the Supplementary Material. Similarly, there are two "$\equiv$" symbols in Equation 5. Why did you choose "$\equiv$"? How are the terms "equivalent" but not "equal"? Do you perhaps mean "$\coloneqq$" for an assignment of variables? That would make more sense to me. "$\equiv$" seems to be an abuse of notation but maybe I am missing something here.
Questions
- How do you explain a decrease in $R_{em}$ and an increase in $D_{em}$ that is caused by (standard) training? - Why did you choose a (wide) ResNet50 and not any other model? Is there a particular reason for that choice? - Why do you think "[...] that trained networks do better than their random counterparts at predicting V1 data for large, but not small, training data set sizes"? Is it possible that the size of the (training) data to fit the regression model is just too small to explain any variance at all or do you have another explanation that rules out that possibility? I think that $N = 135$ is too small to see any effects at all which is why I think there's no notable difference between pretrained and randomly initialized networks in your analyses. However, $N = 3600$ is large enough to observe statistically significant effects. So, I doubt that you can draw any conclusions at all about the smaller of the two datasets. Note that $N = 3600$ is still not a particularly large dataset.
Rating
5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Soundness
2 fair
Presentation
2 fair
Contribution
2 fair
Limitations
Yes. The authors discussed limitations of their work in a dedicated section at the end of the main paper.
Summary
In recent years, DNNs trained on image recognition tasks have emerged as strong predictive models of neural activity in the visual cortex. Typically, a regression model is trained to map DNN responses onto biological neural activity in response to the same inputs. The standard method for evaluating the predictive capacity of given DNN is to quantify the variance explained by the model on held-out images. This compresses the structure of each model into a single scalar value. While this can be used to rank models, it provides little insight into the structure of the model's representations and how they align with the structure of the biological neural representations. Another approach analyzes the similarities in representational geometry between models and neural data. However, the authors, state, these approaches do not relate the geometry of the DNN to the regression model used for the neural predictions. In this paper, the authors leverage a recently proposed method for analyzing the generalization error of a model in terms of the radius and dimensionality of the error along each eigenvector of the input data (images viewed by both model and brain), to analyze the geometry of the generalization error of the regression model that maps a given DNN to neural data. The authors use this method to analyze layer-wise differences in prediction error in a DNN, in terms of the error geometry. They find that the contributions of the error radius and error dimension change across layers. They also analyze difference between trained and untrained DNNs at predicting neural responses. For V1 data, they find that the difference in predictivity and error geometry for trained vs. untrained networks is small. However, for V2 - IT data, they find interesting differences. In particular, they find that training reduces error radius, but increases error dimensionality, i.e. spreads errors more evenly across error modes. They suggest that the improved predictivity of trained networks for V2 - IT data is primarily driven by changes in their intrinsic expressivity rather than significant changes in their representational alignment with neural data. In addition, they examine differences between adversarially vs. standardly trained models. Replicating previous findings, they find that adversarial training overall improves model predictivity. They observe differences in how it affects error mode geometry across layers. In earlier layers, error radius is decreased with little change in error dimensionality. In later layers, error dimensionality is the main distinction between robust and standard models. Lastly, they examine the dependence of the generalization error on the size of the training set of images. They find that for small datasets, the directions with small eigenvalues yield high error. For larger datasets, the additional samples permit learning along small eigenvalue directions.
Strengths
The paper applies a recently proposed methodology for analyzing the generalization error of a model to the problem of understanding DNN-neural data representational alignment. It is a solid idea, and the authors conduct a suite of well-posed experiments that analyze existing results from the literature in this frame.
Weaknesses
The paper would benefit from several changes. First, the motivation of the methodology was unclear. Intuitively, why are we interested in understanding the radius and dimensionality of the error? What insight can this give us into understanding representational alignment? What's the motivation for projecting along the eigenvectors of the data's Gram matrix? I think strong answers to these questions can be formulated. They should appear front and center in the paper. Second, the conclusions obtained by the experiments were difficult to tease out. These should be clearly and simply stated, and presented in a list as an overview of the results. The way the results are written and presented visually, the reader needs to do a lot of digging to parse the meaning and significance. Consider making more intuitive graphical figures to emphasize the main points / contributions / conclusions, and/or condensing the results figures into simpler plots that highlight the important trends. Lastly, the standard approach for interpreting DNN-neural data alignment could be more clearly explained and illustrated, so the reader understands how your approach differs. Consider making a schematic plot that highlights this distinction. Overall, I felt the paper required considerable effort to understand. Given that this work draws from several disciplines and relies on a very recently proposed methodology, I believe the authors should do more work to help the reader understand the motivation and significance of this work.
Questions
I'd like to get a more intuitive handle on the significance of the error modes. I expect the eigenvectors of the gram matrix of the image data would be something like a Fourier basis. In this case, we'd expect large eigenvalues for low frequency bases and small eigenvalues for high frequency. Does this analysis essentially explain how error is distributed across frequency content in the images? If so, how does this give us insight into representational alignment?
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Soundness
3 good
Presentation
2 fair
Contribution
2 fair
Limitations
The significance of the conclusions drawn from the experiments is not clear. The purpose of this tool is to provide insight into neural representations in DNNs and the brain. Greater effort should thus be made to make the meaning of the quantified values (error radius / dimension along eigenvalues of the data's Gram matrix) interpretable.
Thanks for the clarifications
Dear authors, thanks for the clarifications. I have read the rebuttal.
Thank you for the suggestions to improve and strengthen our paper!
I thank the authors for the thorough rebuttal. The rebuttal addressed my concerns, which primarily regard the clarity of the results and presentation of the main ideas. With these clarifications incorporated into the paper, I think it will be a strong and valuable piece of work. I have adjusted my score accordingly.
Thank you for your review and suggestions for how to clarify the presentation and clarity of the paper, and for thinking that this will be a strong piece of work. If you have any further questions we are happy to address them.
Thanks for the thorough response
I thank the authors for taking the time to read my review and respond to it thoroughly. I am glad to see that my feedback was useful. Because I appreciate the effort that the authors made to improve their submission and I feel the responsibility to acknowledge that, I will raise my score from 3 to 5.
Thank you for your efforts to address my concerns. The presentation of this paper is overall improved. I will upgrade my score to 7.
Thank you for the comments and clarification. I will increase my score by 1.
Decision
Accept (spotlight)