Learning Symmetries via Weight-Sharing with Doubly Stochastic Tensors

Group equivariance has emerged as a valuable inductive bias in deep learning, enhancing generalization, data efficiency, and robustness. Classically, group equivariant methods require the groups of interest to be known beforehand, which may not be realistic for real-world data. Additionally, baking in fixed group equivariance may impose overly restrictive constraints on model architecture. This highlights the need for methods that can dynamically discover and apply symmetries as soft constraints. For neural network architectures, equivariance is commonly achieved through group transformations of a canonical weight tensor, resulting in weight sharing over a given group $G$. In this work, we propose to learn such a weight-sharing scheme by defining a collection of learnable doubly stochastic matrices that act as soft permutation matrices on canonical weight tensors, which can take regular group representations as a special case. This yields learnable kernel transformations that are jointly optimized with downstream tasks. We show that when the dataset exhibits strong symmetries, the permutation matrices will converge to regular group representations and our weight-sharing networks effectively become regular group convolutions. Additionally, the flexibility of the method enables it to effectively pick up on partial symmetries.

Paper

Similar papers

Peer review

Reviewer xU9K3/10 · confidence 3/52024-07-10

Summary

In contrast to many works that impose strict architectural constraints to parameterize neural networks that are exactly equivariant to a known group, this work considers learning approximate equivariances to unknown groups. This is done using soft weight sharing schemes, where doubly stochastic tensors are learned and applied on canonical weights. The approach generalizes GCNNs, and is shown to learn natural symmetries in experiments.

Strengths

1. Nice background section. In my opinion, it is written much better than other related papers in the area. 2. The method is a natural generalization of group convolutions, which is developed via the neat perspective of group convolutions as applications of transformations of a canonical weight tensor. It is nice that you can choose the number of "group elements", so that e.g. non-group-symmetries can be captured.

Weaknesses

1. Section on "Regular representaitons allow for element-wise activations" is a little out of place and imcompletely justified. The claim that non-elementwise activations "in practice are not as effective as the classic element-wise activations" is a strong statement that needs more specific justification. 2. Direct parameterization of the kernels for each group element is expensive. Empirical runtime and memory analysis would be appreciated, to see the effect of this. 3. Empirical results are rather limited and weak. Few baselines are considered (see below as well), the baselines already perform well on these tasks, and WSCNN does not improve over the baselines in Table 1. Given that the method does not appear to be too scalable (see above), it is unclear where it would be useful in improving performance. 4. Other baselines (besides standard CNNs and GCNNs) and ablations are missing. For instance, one can imagine removing the sinkhorn projection onto the doubly stochastic matrices. If we initialize the $\Theta_i^l$ to be permutation matrices (your fixing of one representation to be identity is related to this), then this may also perform similarly.

Questions

1. How do the learned filters of WSCNN look on CIFAR10? 2. Do learned representations ever look the same across different "group" elements?

Rating

3

Confidence

3

Soundness

3

Presentation

3

Contribution

2

Limitations

Some discussion in section 6

Reviewer uLBU4/10 · confidence 5/52024-07-11

Summary

Overview: This paper claims to build upon a long line prior work on symmetry detection and learning equivariant kernels. [18,19,20,22, 25-31. Eespecially reference [30] which is an identical weight and parameter sharing scheme that learns and discovers equivariances solely by row-stochastic entries. This paper's idea is to enforce both row and column stochasticity through use of the Sinkhorn operator to achieve equivariance discovery and applicable to more interesting data domains such as images. Advantages: * The paper addresses an important problem of Kernel equivariance via symmetry detection and weight in CNN . However its marginal progress over prior work. Weaknesses: * Comparisons with prior work are missing. * The quantitative advantage of double stochasticity from row and column has advantages for metric spaces (images),. However this is not explained nor quantified through comparison. * Moreover while ground truth tests are done for some toy problems, the quantitative generalizability of the double stochasticity over single stochasticity are never delineated, nor proved. Questions: This paper is identical to this ICML 2024 workshop paper. https://scholar.google.com/citations?view_op=view_citation&hl=en&user=gfRkDXEAAAAJ&citation_for_view=gfRkDXEAAAAJ:ufrVoPGSRksC ? Also why no comparisons to even reference [30] ? Missing References: * Several references are missing, including the entire suite of probabilistic symmetry detection and equivariant NN is not discussed. See for e.g. @misc{bloemreddy2020probabilisticsymmetriesinvariantneural, title={Probabilistic symmetries and invariant neural networks}, author={Benjamin Bloem-Reddy and Yee Whye Teh}, year={2020}, eprint={1901.06082}, archivePrefix={arXiv}, primaryClass={stat.ML}, url={https://arxiv.org/abs/1901.06082}, }

Strengths

This paper enforces both row and column stochasticity through use of the Sinkhorn operator to achieve equivariance discovery. However it's not clear what quantitative benefits this paper's method makes.

Weaknesses

Weaknesses: * Comparisons with prior work are missing. * The quantitative advantage of double stochasticity from row and column has advantages for metric spaces (images),. However this is not explained nor quantified through comparison. * Moreover while ground truth tests are done for some toy problems, the quantitative generalizability of the double stochasticity over single stochasticity are never delineated, nor proved. Questions: This paper is identical to this ICML 2024 workshop paper. https://scholar.google.com/citations?view_op=view_citation&hl=en&user=gfRkDXEAAAAJ&citation_for_view=gfRkDXEAAAAJ:ufrVoPGSRksC ? Also why no comparisons to even reference [30] ? Missing References: * Several references are missing, including the entire suite of probabilistic symmetry detection and equivariant NN is not discussed. See for e.g. @misc{bloemreddy2020probabilisticsymmetriesinvariantneural, title={Probabilistic symmetries and invariant neural networks}, author={Benjamin Bloem-Reddy and Yee Whye Teh}, year={2020}, eprint={1901.06082}, archivePrefix={arXiv}, primaryClass={stat.ML}, url={https://arxiv.org/abs/1901.06082}, }

Questions

Questions: This paper is identical to this ICML 2024 workshop paper. https://scholar.google.com/citations?view_op=view_citation&hl=en&user=gfRkDXEAAAAJ&citation_for_view=gfRkDXEAAAAJ:ufrVoPGSRksC ? Is this ok ?? Also why no comparisons to even reference [30] ? Missing References: * Several references are missing, including the entire suite of probabilistic symmetry detection and equivariant NN is not discussed. See for e.g. @misc{bloemreddy2020probabilisticsymmetriesinvariantneural, title={Probabilistic symmetries and invariant neural networks}, author={Benjamin Bloem-Reddy and Yee Whye Teh}, year={2020}, eprint={1901.06082}, archivePrefix={arXiv}, primaryClass={stat.ML}, url={https://arxiv.org/abs/1901.06082}, }

Rating

4

Confidence

5

Soundness

3

Presentation

2

Contribution

2

Limitations

There are no attempts on analyzing the generalizability of the operator to varying initial and boundary conditions or Reynold numbers. Other potentially relevant paper to consider: https://www.sciencedirect.com/science/article/pii/S0021999123001997 This paper uses error-correction in neural operators based on the residual.

Authorsrebuttal2024-08-12

Dear reviewer uLBU, we believe to have addressed any weaknesses raised in your original review, an acknowledgment or response to our rebuttal would be much appreciated. We are very much open to address any concerns in the remaining discussion period.

Reviewer 3Y5V7/10 · confidence 4/52024-07-12

Summary

The paper proposes a symmetry discovery through learning parameter-sharing in weight matrices. The parameterization relies on relaxing underlying permutation matrices by transforming them as doubly stochastic matrices. In combination with additional regularization, the parameterization can be used to successfully discovery symmetries in data, improving generalization performance.

Strengths

The method appears to be very elegant in offering a natural relaxation of underlying weight-sharing scheme through Sinkhorn operator. The paper is well-written and provides nice illustrations which guide the reader in understanding the proposed method. Then, the paper provides thorough empirical validation demonstrating usefulness of the approach in practice.

Weaknesses

-Regularization. The paper proposes a novel parameterization that can be used for symmetry discovery. In terms of objective, the work relies on direct regularization to avoid equivariants solutions, similar to some other symmetry discovery works. It has been shown in prior work that this strategy can have issues, since it introduces an additional hyperparameter which may need additional tuning (thereby is less ‘automatic’). In terms of explaining the methodology, the paper would benefit from some discussion on the role of the used regularizer. -Scalability. The paper seems to scale quadratically in |X|. Is this not an issue? How does this compare against scaling of alternative methods?

Questions

-Regularization For experiments, it would be helpful if it is clear how the strength of this regularizer is chosen. Cross-validation? -Analysis of learned symmetries It would be interesting to better understand what weight-sharing is being learned for CIFAR-10, apart from merely measuring improved test accuracy. Is there an analysis on the learned symmetry? -Regularization Entropy regularization seems to hurt on CIFAR-10, but improve on MNIST experiments. Do authors have an intuition on why this is the case?

Rating

7

Confidence

4

Soundness

3

Presentation

3

Contribution

4

Limitations

The method proposes an elegant and novel symmetry discovery method. I expect the proposed parameterization to be beneficial to the community. The paper could improve a bit on the objective function side of things (seems to rely on directly encouraging symmetry, like some prior work). Apart from the remaining questions, the paper makes for a strong contribution.

Authorsrebuttal2024-08-12

Dear reviewer 3Y5V, we believe to have addressed any weaknesses raised in your original review, an acknowledgment or response to our rebuttal would be much appreciated. We are very much open to address any concerns in the remaining discussion period.

Reviewer 3Y5V2024-08-12

I appreciate the author's further explanations and comments. I remain my score of a 7 and recommend acceptance, as I deem this a technically solid paper that will have high impact in the community. This is conditional on author's including some more discussion on the use of regularization and the computational concerns, as well as the issues raised by other reviewers.

Authorsrebuttal2024-08-14

We thank you for your response. We are glad you appreciated the work and will incorporate the necessary alterations.

Reviewer 1myY6/10 · confidence 4/52024-07-21

Summary

This paper introduces a parameterisation that contains the ability to represent weight tying corresponding to arbitrary (?) group equivariances. In practice it can represent interpolations of weight tying, but it is argued that this is a feature (not a bug!) since strict equivariance is often too strong a constraint to place on a model for it to still fit the training data. Strict equivariance is known to be represented by a permutation matrix (correct me if I'm wrong), which justifies the need to use doubly stochastic matrices to interpolate. This makes the edges of the parametisation strict equivariances. (For equivariances of discretised continuous signals this is not quite true, as the permutation in continuous space would need to be approximated by some interpolation in the discrete space) This leads to a new parameterisation of a weight structure in a layer that can be trained in the usual way. The doubly-stochastic matrices can then be investigated to see if equivariance is actually learned. The experiments implement the method, and run on benchmark datasets and synthetic datasets, showing good performance, and somewhat interpretable group structure appearing. It is unclear what the actual point of the experiments is, since there are many reasons to use equivariance, but the experiments are not phrased in terms of this (see discussion).

Strengths

The problem of learning equivariances is very important, as it would remove a significant difficulty in designing networks with the correct inductive biases. The solution is flexible, as any (?) group structure can be represented by the parameterisation. There is also an elegant solution to the problem of needing a large number of parameters, that will work in practice for image data: Assuming translational equivariance, and only parameterising additional equivariances on the filters that are much smaller than the image.

Weaknesses

Overall, the paper presents a well-reasoned method to an important problem, and I do believe that it meets the standard for publication at NeurIPS. **Method & Presentation** The method is well-justified. However, a final summary of what a forward pass through a layer looks like was not given, and would be really helpful. In addition, it would be helpful to have a clearer discussion of how many additional parameters are added (beyond lines 220 onwards), with the architectures that are discussed given as an explicit example. Essentially, one thing which seems to be the case, but is not explicitly acknowledged, is that this method collapses to just a special parameterisation of weight matrices, where the weights have low-rank combined with doubly-stochastic structure. The low-rank-ness is shown in eq 6. While this is a simplistic way of looking at the method, it does give a helpful alternative view. Making this explicit would help the paper. **What is the claim of the experiments?** The experiments are the main weakness of the paper. Some qualitative results about the structure of the learned weights are given, which are helpful. But it is not clear what the quantitative claim of the experiments is. Equivariance can help in several ways, e.g. better out-of-distribution prediction, better prediction at low data, or smaller/compacter models. So is the claim that the equivariance inductive bias helps, and it can be discovered automatically? But in this case this is not disentangled from the model capacity. Perhaps a normal CNN would perform better if it were just made larger! This is additionally indicated that the baseline 70% accuracy on CIFAR 10 is low compared to what other non-group-equivariant methods can achieve. This unclarity also exists in the synthetic experiments, where the size of the dataset is not discussed. Low-data experiments could help here, since it's easier to make the model large enough that size doesn't help any longer, which isolates inductive bias only. The CIFAR experiments show that making the model larger improves performance. How can we be sure that this is really the benefit of learning equivariances, rather than adding more capacity? Another experiment that is necessary here, is a comparison to a weight structure that does not have the doubly-stochastic constraint enforced on it. This would allow the effect of simply adding capacity to be tested. Alternatively, a low-data experiment would allow the generalisation capabilities of the model to be tested (e.g. [Invariance Learning in Deep Neural Networks with Differentiable Laplace Approximations](fig 3 in https://proceedings.neurips.cc/paper_files/paper/2022/file/50d005f92a6c5c9646db4b761da676ba-Paper-Conference.pdf)). This could be done on MNIST variants, where currently the differences are so small, it is hard to draw conclusions. This same issue pops up in the synthetic experiments. What is the dataset size? If the dataset is so large that all transformed signals are in the dataset, even a fully-connected network would learn the correct function. A low-data experiment is needed to really show that truly an equivariance has been learned that can help to _generalise_. Alternatively, you need to argue a benefit on the basis of parameter count. In summary: The claims that the experiment section support are not clear. The field of equivariances is mature enough that the potential benefits have been clearly described, and these need to be clearly evaluated in experimental sections. **Related Work** The idea of relaxing equivariance by placing a distribution over transformations is older than the papers currently cited. E.g.: - [Local Group Invariant Representations via Orbit Embeddings](http://proceedings.mlr.press/v54/raj17a/raj17a.pdf) The discussion of "symmetry discovery methods" is unclear. What data do these methods need, or what kind of training signal do they use? What kind of predictive improvements do they obtain? Is the goal of these papers the same as those in the previous paragraph? Or, is the way that these methods learn group structure different from those in the previous section? If so, how? Methods that learn a degree of equivariance on a layer-by-layer basis are relatively new, and it would be good to discuss this explicitly. E.g.: - [Residual Pathway Priors for Soft Equivariance Constraints](https://proceedings.neurips.cc/paper/2021/file/fc394e9935fbd62c8aedc372464e1965-Paper.pdf) Finzi et al is not mentioned at all, but does allow partial equivariance in a layer-by-layer way. - Reference [26] "Learning layer-wise equivariances automatically using gradients" is cited, but only in the context of arguing that overly constrained models suffer poor performance, but not in the context that this paper also discusses how to learn the right equivariance to use. This author has more papers on learning invariance/equivariances that may be relevant. Overall, the literature review misses a lot of relevant work. The suggestions I gave are off the top of my head and I definitely missed some important papers too. However, it is the responsibility of the authors to put the time and effort into going beyond this to give a more thorough overview. **Minor** - _"Requiring no prior knowledge of the possible symmetries."_ (line 52) It is true that earlier methods (with the exception of [31]) could only pick between groups that were completely specified a-priori. While this paper _in principle_ does provide a parameterisation that can _search_ over a much wider space, this space needs to be limited for scalability reasons, and it was not demonstrated that the method would work reliably without this "prior" being added! - It would be really helpful to have a full discussion of the impact on the number of parameters, for these specific experiments.Needs a discussion of the total number of parameters - Is there a typo in params in table 1? Under "Params" should "103 + 265K" be "103K + 256K"?

Questions

- Can this parameterisation represent arbitrary group equivariances? Is the representation of every strict equivariance a permutation matrix? It would be helpful to be explicit about how general this really is. - Am I right in understanding that this method ultimately just parameterises a low-rank weight matrix over different feature channels? - Can the benefit that equivariance promises in the low-data regime still be provided by this method when the invariance is learned? Since effectively, you're just parameterising weights in a different way (low rank?). In rotationally equivariant settings in low-data, could these weights not just overfit, rather than learning to rotate filters? Would this not lose an important benefit of equivariance? - How important is the doubly-stochastic nature of things? Could you just run an experiment without the Sinkhorn component at all? - Can you give a very short (ideally 1 sentence, or a 2-3 sentences) summary of the quantitative claims that are made about the method, that are verified in the experiment section?

Rating

6

Confidence

4

Soundness

2

Presentation

2

Contribution

3

Limitations

See above. Overall, this is a really interesting idea. I do have some concerns about the evaluation. These concerns are large in the scheme of determining how well this method really works relative to clearly formulated claims, but small relative to typical approaches in the ML community.

Authorsrebuttal2024-08-12

Dear reviewer 1myY, we believe to have addressed any weaknesses raised in your original review, an acknowledgment or response to our rebuttal would be much appreciated. We are very much open to address any concerns in the remaining discussion period.

Reviewer 1myY2024-08-14

> We acknowledge that conventional CNNs with an equivalent channel count inherently possess a higher expressive capacity due to a lack of kernel constraints. Therefore, we have opted to compare overall parameter counts rather than the number of kernels, a more comprehensive assessment in our opinion. For completeness, we have additionally measured the performance for an unconstrained 128-channel CNN (please consider point 4 in the general rebuttal). This is a really helpful addition. Interesting, so the main claim now is that for a parameter-constrained model, you obtain better performance? Size of the model clearly does come into this, as the CNN also gets better as more filters are added. This is fine, and a clear claim. However this does mean that the current experiments only show that the invariance learning properties work in the _underfitting_ regime! This is different from what is wanted from learning equivariances, where we may want to show the ability to obtain improved performance, once other simpler methods like making the model larger stop working. This is why other methods e.g. [31, 26] consider second order information (meta-learning, marginal likelihood approximations): to distinguish between inductive biases even when the training loss cannot. This was also discussed in [*] and back in 2018 [**]. This is what the low data experiment could have shown, if it was verified that the model was large enough to sufficiently fit the training data. Also the low-data experiment is of limited usefulness, since it does not contain a baseline of a non-invariant model. The interesting question here is to see how much of an improvement you can get by learning invariances over a model that cannot do this. (Also relevant to your answer to my question 4.) > We would like to stress that we do not explicitly impose any group structure in our representation stack This was clear from the paper, and is certainly valuable! However, I don't see how this makes the literature I pointed to any less relevant? ## Overall Either way, while I don't think this is the end of the question of how to learn equivariance automatically, I do believe it is an interesting paper, and I will continue to argue for acceptance. [*] [Invariance Learning in Deep Neural Networks with Differentiable Laplace Approximations](https://proceedings.neurips.cc/paper_files/paper/2022/file/50d005f92a6c5c9646db4b761da676ba-Paper-Conference.pdf) 2022 (not cited in the paper) [**] [Learning Invariances using the Marginal Likelihood](https://arxiv.org/abs/1808.05563) 2018 (not cited in the paper)

Area Chair ijCz2024-08-11

Reviewer-Author discussion period

Dear Reviewers, The deadline for the reviewer-author discussion period is approaching. If you haven't done so already, please review the rebuttal and provide your response at your earliest convenience. Best wishes, AC

Reviewer xU9K2024-08-12

Hello authors. Thank you for the response, and apologies for my late response. Your elaboration on (Question 1) is helpful (although it is clear from your paper, I forgot since I was thinking in terms of equivariant networks). I think that this, as well as your new ablations on less constraining of the learned representations, would be useful additions to your paper. My remaining worries are more high-level, on the utility of the method. In my view, the utility of approximate equivariance or discovering symmetries is not adequately demonstrated or achieved with the given method. The experiments are only on augmentations of MNIST, or on non-SOTA regimes in CIFAR10. This could be fine, if the model were developed with methods that one could see being useful in different regimes in the future. However, I think the poor efficiency and scalability of the model severely harms this. On the axes of efficiency and accuracy, the model is not efficient, and does not have clear accuracy gains in interesting experimental areas. Perhaps this work would benefit from expriements in areas (say, in the physical sciences) that more clearly desire equivariance, where equivariant models are actually SOTA or near SOTA.

Authorsrebuttal2024-08-12

Thank you for your response, and we are glad some of your concerns have been addressed. Regarding your remaining concern about state-of-the-art performance, it is important to highlight that the primary aim of our approach is in line with regular group convolutional methods, focusing on learning useful constrained functions rather than competing directly with SOTA image models. This distinction is crucial as it frames our method as a tool for understanding and utilizing inherent symmetries in data. Note that the method can easily be extended to any task where signals act on a discretized domain, such as volumetric grids. In these cases there is some evidence that having unconstrained models are more prone to overfitting [1] and our method's ability to inherently recognize and utilize symmetries can be particularly advantageous. For the sake of developing the method, we consider them out of scope for this work. As such, in the current experiments we utilize MNIST and CIFAR-10 to demonstrate the main claims about symmetry discovery and feasibility. Additionally, while many CNNs rely on extensive data augmentation to achieve SOTA results, there is emerging evidence suggesting that such strategies might not always align well with real-world data distributions, potentially leading to suboptimal generalization [2, 3, 4, 5]. By contrast, our constrained approach offers a systematic way to learn and adapt to symmetries present in the data. In direct comparisons, our model shows that it can match the performance of unconstrained CNNs, underlining its efficacy even without clear accuracy gains in the standard experimental setups used (see point 4 in the general response). In these cases the CNN models have a considerably larger number of trainable parameters, implying that our method may not need to operate at the same scale/channel capacity to achieve satisfactory performance. This underlines the method's potential utility in settings where understanding and incorporating data-driven symmetries are crucial. We believe that further development and application of this method in contexts where equivariance is highly valued will substantiate its utility and is left as future work. We kindly encourage considering the broader context of our research’s objectives and its potential contributions to the field. **References** [1] Regular SE(3) Group Convolutions for Volumetric Medical Image Analysis. Thijs P. Kuipers and Erik J. Bekkers. MICCAI 2023 [2] Learning Equivariances and Partial Equivariances from Data. David W Romero, Suhas Lohit. NeurIPS 2022 [3] Learning Invariances in Neural Networks. Gregory Benton, Marc Finzi, Pavel Izmailov, Andrew Gordon Wilson. NeurIPS 2020 [4] Learning Layer-wise Equivariances Automatically using Gradients. Tycho F. A. van der Ouderaa and Alexander Immer and Mark van der Wilk. 2023 [5] Relaxing Equivariance Constraints with Non-stationary Continuous Filters. Tycho FA van der Ouderaa, David W Romero, Mark van der Wilk. NeurIPS 2022

Reviewer xU9K2024-08-13

We thank the authors for sharing their thoughts on the scope and context of the work within the literature. My worries still hold, so I maintain my score. GCNNs came out 8 years ago, and basic unconstrained CNNs on MNIST and CIFAR10 are quite far from empirical practice. In my opinion, experiments that actually significantly benefit from equivariance / approximate equivariance / equivariance detection would substantially improve the paper. There are many tasks / application areas that the community claims do or may benefit from equivariance, so it would be good to see experiments in such areas. I worry that factors such as the scalability of your method could harm its applicability in these tasks / applications, so experiments would be needed to reject this hypothesis.

Authorsrebuttal2024-08-14

Thank you for your continued engagement and feedback on our submission. We appreciate your concerns and take them seriously as they help us refine our approach and clarify our contributions. We understand that the empirical practice has evolved since the introduction of GCNNs, among other models with known/built-in equivariance structures. However, the field of automatic (partial) symmetry discovery is fairly new, and the ability to detect and utilize equivariances automatically and the state of current research (i.e. [1, 2, 3, 4, 5]) operates in controlled and well-understood environments in order to validate the fundamental aspects of the proposed method. While we recognize the desire for experiments in application areas that could directly benefit from equivariance, our current focus is on establishing a strong theoretical and empirical foundation for our method. Such foundational work is essential for understanding the potential and limitations of new techniques before they can be successfully applied to more demanding domains. Your point on scalability and applicability in more practical scenarios is well-taken. We see this as an important direction for future research, where the scalability of our method can be further developed and tested in more complex environments. We plan to explore this in future work. **References** [1] Learning Equivariances and Partial Equivariances from Data. David W Romero, Suhas Lohit. NeurIPS 2022 [2] Learning Invariances in Neural Networks. Gregory Benton, Marc Finzi, Pavel Izmailov, Andrew Gordon Wilson. NeurIPS 2020 [3] Learning Layer-wise Equivariances Automatically using Gradients. Tycho F. A. van der Ouderaa, Alexander Immer, Mark van der Wilk. 2023 [4] Relaxing Equivariance Constraints with Non-stationary Continuous Filters. Tycho FA van der Ouderaa, David W Romero, Mark van der Wilk. NeurIPS 2022 [5] Bispectral Neural Networks, Sophia Sanborn, Christian Shewmake, Bruno Olshausen, Christopher Hillar. ICLR 2023

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC