D-CIPHER: Discovery of Closed-form Partial Differential Equations

Closed-form differential equations, including partial differential equations and higher-order ordinary differential equations, are one of the most important tools used by scientists to model and better understand natural phenomena. Discovering these equations directly from data is challenging because it requires modeling relationships between various derivatives that are not observed in the data (equation-data mismatch) and it involves searching across a huge space of possible equations. Current approaches make strong assumptions about the form of the equation and thus fail to discover many well-known systems. Moreover, many of them resolve the equation-data mismatch by estimating the derivatives, which makes them inadequate for noisy and infrequently sampled systems. To this end, we propose D-CIPHER, which is robust to measurement artifacts and can uncover a new and very general class of differential equations. We further design a novel optimization procedure, CoLLie, to help D-CIPHER search through this class efficiently. Finally, we demonstrate empirically that it can discover many well-known equations that are beyond the capabilities of current methods.

Paper

Similar papers

Peer review

Reviewer ubhm6/10 · confidence 3/52023-07-06

Summary

The paper falls in the realm of data driven discovery of dynamical systems, PDEs to be specific. It proposes a framework for a class of PDEs which are termed as variation ready PDEs. These are claimed to be less restrictive than existing methods, which make stronger assumptions on the form of the PDE to be discovered and hence can't recover a significant population. A new optimization scheme, novel loss function and empirical evidence is provided to support the claims made in the paper.

Strengths

In general the paper is well written, the scope of the paper is well thought. Notations are clear and introduced properly. It address an important problem, the clear listing of challenges in the introduction is particularly impressive. The variational loss function is novel

Weaknesses

My major concern is regarding the way the problem is setup before the solution is proposed. Some of the terms introduced here although intuitive lack enough insight to make them more convincing. Section 7 is too short in the main paper to determine any novelty in the optimization scheme, this definitely needs to be presented better in the main paper, as this is claimed as a contribution. One of the aspects mentioned in the paper is that of robustness from noisy or infrequent observations, I don't have see any evidence to support such claims.

Questions

1) Can authors discuss the effect of the choice of test functions as B-splines? What about other options, how do they impact the results. A discussion on this ground would be helpful. 2) Can authors shed light on the robustness of the proposed framework with respect to noisy and/or infrequent observations. These if theoretical can be taking into account sampling frequency, SNR, etc. 3) How about parameterizing the test functions and making them learnable? Has this been tried, if not what do the authors feel in this regard. While test functions are fairly restrictive, there might still be a way to learn then or at the very least choosing them from a dictionary.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

I don't see direct potential negative social impact. However, since this is a paper in the direction of ML for science and adjacent domains, it is quite possible that these methods when fully developed will have major impact. It will be good to have a word of caution regarding this, to make sure we acknowledge the vast amount of domain expertise already available in all scientific fields and not use ML as a tool to replace all human knowledge. Authors have acknowledged technical limitations appropriately.

Reviewer Znhz6/10 · confidence 3/52023-07-16

Summary

This paper proposes a framework (D-CIPHER) to discover closed-form PDEs and ODEs. The framework is more general than some of the previously existing methods, and in particular can handle a class of PDEs defined as variational-ready PDEs in the paper. The empirical experiments evaluated the discovery performance on synthetic data for a set of different equations (showing both comparisons with existing methods on discovering linear combinations and results on discovering more challenging equations).

Strengths

Originality: From what I could tell, this paper proposes an original framework to discover broader classes of PDEs and ODEs. That being said, I'm not familiar with the literature in the area so I cannot fully speak to the originality. Quality: The empirical experiments, even though synthetic, appear to be thoughtfully designed and demonstrate notable improvements. Clarity: This paper is very well-written with a good balance of technical details and general introductions. I enjoyed reading it even as someone outside of the ODE/PDE field.

Weaknesses

Significance: This is probably my biggest question for the paper. What types of real-world scenarios could D-CIPHER be applied to? The Discussion section briefly mentions finding heat and vibration sources and discovering population models and epidemiological models. However, at the level of the current discussion, these applications all sound very abstract. The paper would benefit from more grounding in concrete applications. If it's possible to add in experiments on real data, that would strengthen the paper. Even if not, the paper would still be improved with a more extensive discussion on how D-CIPHER could be applied in each of the relevant scenarios. For example, what would the form be (in step 1) based on prior knowledge? How would the fields be estimate (step 2)?

Questions

See "weaknesses"

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

4 excellent

Contribution

2 fair

Limitations

The very last paragraph discusses some potential limitations, but in my opinion it would be very helpful to have some "negative examples" in the paper, i.e. synthetic experiments where D-CIPHER fails and to explore the reasons of failure.

Reviewer kC9e7/10 · confidence 2/52023-07-24

Summary

The paper proposes a new way of discovering closed-form Partial Differential Equations (PDEs) from data. This especially aims at high-order PDEs, especially when the specific form is not pre-assumed and there is a lack of observations on derivatives. The key idea is to represent the unknown PDE with terms that are bounded by derivatives and terms that are not, so that the latter kind can be easily and reliably estimated from data, hence the ground-truth, while the first kind can be estimated by leveraging symbolic regression. Several synthetic datasets are employed for evaluation, many of which are simulated from equations that do not satisfy the linear combination assumption made by existing work.

Strengths

Strengths: 1. A new framework for discovering PDEs, especially the ones with high-order derivatives and a lack of direct observations on the derivatives. 2. A comprehensive evaluation on many different PDEs satisfying and beyond the assumption of PDEs made by previous work. 3. The comparison seems to show a better performance. 4. Good exposition. The paper has a good balance between background and technical details.

Weaknesses

I am in general in favour of this paper. However, it is not my area of expertise, so it would be good if they authors could clarify some questions here. 1. Reliance on symbolic regression. I wonder to what extent the proposed framework has to rely on symbolic regression. This opens up several questions. (1) How inclusive or comprehensive does the dictionary has to be? What if some key derivatives are not present in the dictionary? (2) can the authors provide more details on the computational time in addition to E.2? It would be good to show the computation time in F.3, when the dictionary is gradually increased 2. Comparison. It is arguably true that derivatives are hard to measure in applications. But since the experiments are done using synthetic data, I wonder what if the observations of derivatives are available? Will the proposed method still outperform existing methods? I would assume the proposed framework still works to some extent, although not being able to make use of the observed derivatives.

Questions

Please see my 'Weaknesses' section above.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

The authors mentioned limitations. But in real-world scenario, there are other factors that might make applying this framework difficult. The first one is the sparsity of observations. The sensors are normally not well distributed and sometimes extremely sparse. So the estimate based on the zero-order information might not be reliable to start with. The second is the type of noise, which is normally unknown and needs to be estimated with the data together.

Reviewer kC9e2023-08-17

Rebuttal clarifies my questions

Thanks for the detailed responses and the added content. I will keep my confidence low as this is not my area of expertise but I am happy to see this paper accepted.

Authorsrebuttal2023-08-19

Dear Reviewer kC9e, We appreciate your time invested in assessing our paper and the rebuttal. We are delighted to see that you reaffirmed your acceptance of our work! Your constructive comments have played a significant role in enhancing the quality of our paper. Kind regards, Authors of Submission13865

Reviewer ubhm2023-08-18

Thank you for the detailed response, and proposed update. I will keep my score.

Authorsrebuttal2023-08-19

Dear Reviewer ubhm, We want to express our sincere gratitude for the time and effort you dedicated to evaluating our paper and the rebuttal. We are pleased by your continued positive assessment of our work. Your insightful comments have undeniably contributed to the enhancement of our paper. Kind regards, Authors of Submission13865

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC