Can neural operators always be continuously discretized?

We consider the problem of discretization of neural operators between Hilbert spaces in a general framework including skip connections. We focus on bijective neural operators through the lens of diffeomorphisms in infinite dimensions. Framed using category theory, we give a no-go theorem that shows that diffeomorphisms between Hilbert spaces or Hilbert manifolds may not admit any continuous approximations by diffeomorphisms on finite-dimensional spaces, even if the approximations are nonlinear. The natural way out is the introduction of strongly monotone diffeomorphisms and layerwise strongly monotone neural operators which have continuous approximations by strongly monotone diffeomorphisms on finite-dimensional spaces. For these, one can guarantee discretization invariance, while ensuring that finite-dimensional approximations converge not only as sequences of functions, but that their representations converge in a suitable sense as well. Finally, we show that bilipschitz neural operators may always be written in the form of an alternating composition of strongly monotone neural operators, plus a simple isometry. Thus we realize a rigorous platform for discretization of a generalization of a neural operator. We also show that neural operators of this type may be approximated through the composition of finite-rank residual neural operators, where each block is strongly monotone, and may be inverted locally via iteration. We conclude by providing a quantitative approximation result for the discretization of general bilipschitz neural operators.

Paper

References (54)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer 47x86/10 · confidence 2/52024-06-17

Summary

This paper studies the question of whether a neural operator (or a general diffeomorphism on an infinite dimensional Hilbert space) can be continuously discretized through the lens of category theory. It first proves that there does not exist a continuous approximation scheme for all diffeomorphisms. Then, it shows that neural operators with strongly monotone layers can be continuously discretized, followed by a proof that a biliptschitz neural operator can be approximated by a deep one with strongly monotone layers. Some further consequences are discussed.

Strengths

1. The paper studies the discretization of a continuous operator, which is a source of error and an important issue in operator learning if one is not careful. 2. The paper is fairly comprehensive, encompassing both positive and negative results. The study of positive results contains theorems of different flavors. 3. Although the theory is based on the category theory, the presentation and explanation of the results are relatively clear and accessible for people who are unfamiliar with it.

Weaknesses

1. Most results presented in the paper are purely theoretical and lack a quantitative or asymptotic estimate. For example, * the notion of continuous approximation functor in Definition 8 does not care about the rate of convergence, and * there is no estimate for the number of layers $J$ in Theorem 4. 2. No empirical results are there to support the theory. While this is a theoretical paper, some toy experiments that exemplify the theory would be very helpful.

Questions

1. In general, is there any assumption of the Hilbert space studied in this paper? For example, are the Hilbert spaces assumed to be separable? 2. What is the definition of the convergence of finite-dimensional subspaces used in this paper? Strong convergence of the projection operator? I do not think there is a standard definition for this in elementary functional analysis so it would be helpful to say it explicitly in the paper. 3. Have you studied the role of the bilipschitz constant in your theorems? For example, how does the number of layers $J$ in Theorem 4 depend on it?

Rating

6

Confidence

2

Soundness

3

Presentation

3

Contribution

3

Limitations

None.

Reviewer KEed2024-08-13

Well, I felt sorry for the authors because a reviewer of NeurIPS, once well-known for a top theory-oriented conference of machine leaning, **cannot raise his/her score** simply because he/she is **not familiar with the category theory**. I even feel this is symbolic of the current state of a theory-oriented machine learning conference. It should not be the problem of the individual reviewer him/herself, but the problem of conference's matching systems that mistakenly assign a perfect amateur of the category theory for reviewing a category theoretic research. This is a suggestion for chairs for future avoidance of mismatches, that the reviewers should be examined if they have fundamental knowledge/background/understandings in the field. I am an expert of expressive power analysis, but not at all of category theory and tropical geometry. Unfortunately, this kind of mismatch happens every year, so I am usually skeptical to any mathematical ''theorems'' published in machine learning conferences.

Reviewer K5mB7/10 · confidence 4/52024-06-26

Summary

This paper focuses on the continuous discretization in operator learning. This is a very important question since it involves reducing the infinite-dimensional space to a finite-dimensional space in operator learning. The authors present cases where discretization is continuous and cases where it is not. The results are interesting and can be applied to design methods in operator learning.

Strengths

The proof is solid, and the paper is well-written and organized. I appreciate the results presented in this paper.

Weaknesses

Since this paper is submitted to NeurIPS and not a mathematical journal, I hope the authors can provide some practical examples, such as solving the Poisson equation \(\Delta u = f\) to learn the operator relationship between \(f\) and \(u\). By using methods like DeepONet and FNO, it would be beneficial to determine whether the discretization in these methods is continuous or not. I believe this could make the paper more accessible to a broader audience.

Questions

Mentioned in the Weakness.

Rating

7

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

All right.

Reviewer KEed6/10 · confidence 3/52024-07-12

Summary

This paper investigates theoretical limitations of discretizing neural operators on infinite-dimensional Hilbert spaces. The authors first prove a "no-go theorem" (Theorems 1,2) showing that diffeomorphisms between infinite-dimensional Hilbert spaces cannot generally be continuously approximated by finite-dimensional diffeomorphisms. Then, they provide positive results for certain classes of operators such as strongly monotone (Theorem 3) and bilipschitz neural operators (Theorem 4). They finally provide concrete example of approximation by finite residual ReLU networks (Theorem 5).

Strengths

- The universality of neural networks has been demonstrated in various settings. However, research on the approximation abilities of operators is relatively scarce. Particularly, the characterization of classes that cannot be approximated is intriguing. This study is important as it succinctly demonstrates the differences between finite-dimensional and infinite-dimensional properties in the manageable setting of Hilbert spaces. - Moreover, the novel approach of expressing approximation sequences in terms of category theory is noteworthy.

Weaknesses

- On the other hand, the proofs are based on conventional analytical arguments rather than category-theoretic arguments. Therefore, the "category theory" framework might be somewhat exaggerated. It is expected that with refinement of notation and sentence structure, the description could become more perspicuous in the future. - There is concern that the categorical description may have obscured the contributions typically seen in __traditional approximation theory__ papers. As the authors likely recognize, various topologies are used in function approximation, and this study focuses __only__ on approximation in the norm topology of Hilbert spaces, and does not negate "all considerable approximation sequences". So, the impossibility theorem presented here might simply be due to the norm topology being too strong. While the Hilbert structure sounds natural as a generalization of Euclidean structure, in reality, concepts like L2 convergence of Fourier series are quite technical and not necessarily an inevitable notion of convergence. It seems that in pursuit of an elegant categorical description, the diversity of function approximation may have been compromised.

Questions

In Definition 3, why $\sigma$ is imposed besides $G$?

Rating

6

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

The authors did not discuss the validity of assumptions.

Reviewer mJde4/10 · confidence 2/52024-07-15

Summary

The paper addresses the problem of discretizing neural operators, maps between infinite dimensional Hilbert spaces that are trained on finite-dimensional discretizations. Using tools from category theory, the authors provide a no-go theorem showing that diffeomorfisms between Hilbert spaces may not admit continuous approximations by diffeomorfisms on finite spaces. This highlights the fundamental differences between infinite-dimensional Hilbert spaces and finite-dimensional vector spaces. Despite these challenges, the authors provide positive results, showing that strongly monotone diffeomorphism operators can be approximated in finite dimensions and that bilipschitz neural operators can be decomposed into strongly monotone operators and invertible linear maps. Finally, they observe how such operators can be locally inverted through an iteration scheme.

Strengths

- The paper provides theoretical results addressing the challenging problem of discretizing inherently infinite-dimensional objects (neural operators)

Weaknesses

- The text and presentation require significant polishing. It contains numerous typos, poorly formulated sentences, and instances of missing or repeated words - While the paper's theoretical focus is valuable, it lacks examples of specific neural operator structures that meet the theorems or remarks - A more detailed discussion on the practical impact of this work, accompanied by examples, would be beneficial for the audience Please note that my review should be taken with caution, as I am not familiar with category theory and did not thoroughly check the mathematical details. My feedback primarily focuses on the presentation and potential impact of the results rather than a rigorous validation of the theoretical content.

Questions

- Neural operators are typically defined between Banach spaces. Why does your theory focus on maps between Hilbert spaces instead? - Comment: The work in [1] might have been relevant to cite as well. - The main neural operator paper [2] develops theoretical results on the universal approximation theory of neural operators. How do your results relate to the ones in that paper? [1] F. Bartolucci, E. de Bézenac, B. Raonić, R. Molinaro, S. Mishra, R. Alaifari, Representation Equivalent Neural Operators: a Framework for Alias-free Operator Learning, NeurIPS 2023. [2] Nikola B. Kovachki, Zongyi Li, Burigede Liu, Kamyar Azizzadenesheli, Kaushik Bhattacharya, Andrew M. Stuart, and Anima Anandkumar. "Neural operator: Learning maps between function spaces with applications to PDEs," J. Mach. Learn. Res., 24(89):1–97, 2023.

Rating

4

Confidence

2

Soundness

2

Presentation

1

Contribution

2

Limitations

The paper lacks examples of applications of the theorems to specific neural operator structures, and some further discussion on the practical impact of the results with examples

Reviewer K5mB2024-08-07

Thanks for the reply. I will keep my score.

Reviewer KEed2024-08-12

Thank you for detailed clarifications. I would like to keep my score as is. > The proofs are indeed based on analytical arguments. If so, I recommend the authors to reconsider the following phrases in the abstract and conclusion: > Using category theory, we give a no-go theorem > We used tools from category theory to produce a no-go theorem It would be much impactful and significant if the authors could more directly point out any incorrectness of the proof or inappropriateness of the assumption in the previous studies.

Authorsrebuttal2024-08-13

Thank you for your response. We would will address your points in the following way. - If so, I recommend the authors to reconsider the following phrases in the abstract and conclusion: "Using category theory, we give a no-go theorem." "We used tools from category theory to produce a no-go theorem" We appreciate the advice, and will follow it. We will replace "Using category theory, we give a no-go theorem" with "Using analytical arguments, we give a no-go theorem framed with category theory." and replace "We used tools from category theory to produce a no-go theorem" with "We give a no-go theorem framed with category theory" in the abstract and conclusion. - It would be much impactful and significant if the authors could more directly point out any incorrectness of the proof or inappropriateness of the assumption in the previous studies. There are several papers which use continuous functions (either as elements of infinite dimensional function spaces or metric spaces) to model images or signal and apply statistical methods and invertible neural networks or maps modeling diffeomorphisms. Often in these papers one derives theoretical results in the continuous models and presents numerical results using a finite dimensional approximations. In this process the errors are caused by the discretization and the effect of changing the dimension of the approximate models. We believe that our work meaningfully addresses these questions as applied to injective/bijective neural operators, an important architecture. We hope that our paper inspires further study these points. We can include citations to the following papers, related to these issues. The below papers which combine neural networks and approximation of diffeomorphisms, as applied to imaging. - Elena Celledoni · Helge Glöckner · Jørgen N. Riseth, Alexander Schmeding Deep neural networks on diffeomorphism groups for optimal shape reparametrization. BIT Numerical Mathematics (2023) 63:50 - GradICON: Approximate Diffeomorphisms via Gradient Inverse Consistency Lin Tian · Hastings Greer · François-Xavier Vialard · Roland Kwitt · Raúl San José Estépar · Richard Jarrett Rushmore · Nikolaos Makris · Sylvain Bouix · Marc Niethammer West Building Exhibit Halls ABC 153 The below papers combine invertible neural networks and statistical models, especially for solving inverse problems (including imaging problems). - Alexander Denker , Maximilian Schmidt , Johannes Leuschner and Peter Maass Conditional Invertible Neural Networks for Medical Imaging. Journal of Imaging 2021, 7(11), 243 - Ardizzone, L.; Kruse, J.; Rother, C.; Köthe, U. Analyzing Inverse Problems with Invertible Neural Networks. In Proceedings of the 7th International Conference on Learning Representations (ICLR 2019), New Orleans, LA, USA, 6–9 May 2019. - Anantha Padmanabha, G.; Zabaras, N. Solving inverse problems using conditional invertible neural networks. J. Comput. Phys. 2021, 433, 110194 - Denker, A.; Schmidt, M.; Leuschner, J.; Maass, P.; Behrmann, J. Conditional Normalizing Flows for Low-Dose Computed Tomography Image Reconstruction. In Proceedings of the ICML Workshop on Invertible Neural Networks, Normalizing Flows, and Explicit Likelihood Models, Vienna, Austria, 18 July 2020. - Hagemann, P.; Hertrich, J.; Steidl, G. Stochastic Normalizing Flows for Inverse Problems: A Markov Chains Viewpoint. SIAM/ASA Journal on Uncertainty QuantificationVol. 10, Iss. 3 (2022) 10.1137 - Papamakarios, G.; Nalisnick, E.T.; Rezende, D.J.; Mohamed, S.; Lakshminarayanan, B. Normalizing Flows for Probabilistic Modeling and Inference. Journal of Machine Learning Research 22 (2021) 1-64

Reviewer 47x82024-08-12

Thank you for the detailed response. Since I am not absolutely familiar with category theory and other related work, I am unable to further raise my score, but I acknowledge that I have read through the rebuttal and it appears to be a nice paper overall.

Reviewer mJde2024-08-13

Thank you for your detailed response. As I mentioned in my initial review, my understanding of category theory is somewhat limited. My feedback has mainly focused on the presentation and potential impact of the results rather than an in-depth validation of the theoretical content. I am not in a position to increase my score.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC