Equivariant Neural Diffusion for Molecule Generation

We introduce Equivariant Neural Diffusion (END), a novel diffusion model for molecule generation in 3D that is equivariant to Euclidean transformations. Compared to current state-of-the-art equivariant diffusion models, the key innovation in END lies in its learnable forward process for enhanced generative modelling. Rather than pre-specified, the forward process is parameterized through a time- and data-dependent transformation that is equivariant to rigid transformations. Through a series of experiments on standard molecule generation benchmarks, we demonstrate the competitive performance of END compared to several strong baselines for both unconditional and conditional generation.

Paper

References (56)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer 5dRq6/10 · confidence 3/52024-07-05

Summary

This paper presents Equivariant Neural Diffusion (END), a novel diffusion model for 3D molecule generation. The major novelty of END over previous molecule diffusion models lies in adopting a learnable forward process based on neural flow diffusion models. Experiments show that END can achieve good performance on benchmark datasets.

Strengths

- Successfully incorporate neural flow diffusion models into the equivariant 3D molecule generation framework and demonstrate its usefulness in experiments. - Generally good, clear and well-organized writing.

Weaknesses

- The novelty contribution of the proposed END method is not high, as END is largely a combination of EDM [1] and neural flow diffusion models. Particularly, the paper does not give a clear discussion or analysis about why adopting learnable forward diffusion process is useful and beneficial to 3D molecule generation, or what molecular structures can be additionally captured by END through learnable forward diffusion process compared with previous diffusion models. - There already exist some SDE based 3D molecule generation methods like EDM-BRIDGE [2] and EEGSDE [3]. Authors are encouraged to highlight the key difference in the diffusion process between END and these methods. - Compared with GEOLDM, END does not show better performance (Table 1 and 2), which weakens the claim about the advantages of using learnable forward process. Since the main evaluation metrics proposed by EDM [1] in 3D molecule generation are saturating in recent literatures, authors are encouraged to adopt metrics proposed by HierDiff [4] to better evaluate the quality of generated 3D molecules. [1] Equivariant Diffusion for Molecule Generation in 3D. ICML 2022. [2] Diffusion-based Molecule Generation with Informative Prior Bridges. NeurIPS 2022. [3] Equivariant Energy-Guided SDE for Inverse Molecular Design. ICLR 2023. [4] Coarse-to-Fine: a Hierarchical Diffusion Model for Molecule Generation in 3D. ICML 2023. -------------------Post Rebuttal--------------------- I appreciate authors' efforts in addressing my concerns and questions in rebuttal. After reading over authors' rebuttal responses and pdf, I think all my concerns have been addressed so I increased my score. I hope authors will carefully add all rebuttal updates (discussion, analysis and experiment results) to the revised version of paper in the future.

Questions

See Weaknesses part.

Rating

6

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

Yes.

Reviewer jNQ26/10 · confidence 3/52024-07-13

Summary

Paper presents, END, a diffusion models for 3D molecule generation that - is equivariant to euclidean transformations and - includes a learnable forward process. Specifically, the forward process in the presented models, is defined as a learnable transformation, dependent on both time and data such that the resulting latent representation $z_t$ transforms covariantly with the injected noise. This is the main difference between the forward pass of Neural Flow Diffusion Models and the proposed method.

Strengths

- Can be used for both conditional and conditional molecules generation. - Improves on existing equivariant diffusion models. - Experimentally, the proposed method shows improvements on both conditional and unconditional generation on the QM9 and GEOM-DRUGS datasets.

Weaknesses

- Lack of a thorough ablation of the proposed method.

Questions

- The only form of ablation i see for the proposed model is in Table 1 where two versions of END are provided. However even here, it seems the END with $\mu_\phi$ performs similar or sometimes better (according to the presented metrics) than the full END model. Is the same pattern observed for the conditional generation tasks? - Relating to the first question above, can a much more thorough ablation of the proposed method be provided in both testing scenarios to really ascertain the utility of the components in the proposed method? - Can the training times for the benchmarked methods be provided for comparison? While it is stated passingly that END requires more training time, can this be cast in contrast with the baselines by providing the actual numbers?

Rating

6

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

Limitations are adequately discussed by the authors.

Reviewer K86T7/10 · confidence 4/52024-07-13

Summary

The Equivariant Neural Diffusion (END) model is a novel approach for molecule generation in 3D that maintains equivariance to Euclidean transformations. Unlike traditional diffusion models that use a pre-specified forward process, END introduces a learnable forward process, parameterized through a time- and data-dependent transformation. This innovation allows the model to adapt better to the underlying data distribution. Experimental results demonstrate that END outperforms strong baselines on standard benchmarks for both unconditional and conditional generation tasks, particularly excelling in generating molecules with specific compositions and substructures. This flexibility in modeling complex molecular structures suggests significant potential for applications in drug discovery and materials design.

Strengths

**Originality:** END introduces a novel learnable forward process, diverging from the fixed processes in traditional diffusion models. This allows the model to better adapt to the underlying data distribution, especially in complex 3D molecular structures. It creatively combines elements from Neural Function Matching Diffusion Models (NFDM) and Equivariant Diffusion Models (EDM), enhancing the generative process while maintaining E(3) equivariance. **Quality:** The model demonstrates superior performance in generating 3D molecular structures, outperforming strong baselines in both unconditional and conditional settings. It excels in generating stable, valid, and unique molecules, particularly evident in its results on the GEOM-DRUGS dataset. Comprehensive experiments and ablation studies confirm the robustness and reliability of the model. **Clarity:** The paper provides a detailed and clear exposition of the methodology, including the formulation of the learnable forward process, parameterization, and evaluation metrics. The inclusion of algorithmic steps and extensive experimental details facilitates replicability for researchers. **Significance:** END represents a substantial advancement in generative modeling for 3D molecules, addressing limitations of prior models by improving sample quality and generation speed. Its ability to maintain equivariance while achieving high performance has significant implications for applications in drug discovery and materials design, potentially transforming these fields.

Weaknesses

**Performance Consistency:** The performance of END is not consistently superior to existing baselines. Although it shows competitive results, there are instances where traditional models, like EDM and its variants, outperform END, particularly in metrics such as validity and uniqueness across different datasets. **Complexity and Scalability:** The added complexity of a learnable forward process, while innovative, increases the model’s training time and resource requirements. END requires more computational resources and longer training periods compared to simpler, fixed-process models like EDM . The model operates on fully-connected graphs, which limits its scalability to larger datasets and more complex molecular structures. This constraint can hinder its applicability in more demanding real-world scenarios.

Questions

- Why was GeoDiff not included as a baseline in your comparisons? - Could you provide more detailed ablation studies that isolate the contributions of the learnable forward process and other key components of END?

Rating

7

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

**Computational Complexity:** The learnable forward process in the END model increases computational complexity, resulting in longer training times and higher resource usage. The authors should discuss potential optimizations to reduce this overhead, such as more efficient algorithms or hybrid approaches. Comparing training times and resource requirements with simpler models would provide useful insights into the model's efficiency. **Scalability Issues:** END's architecture limits its scalability to larger datasets and more complex molecular structures like proteins. To enhance scalability, the authors should explore strategies like sparse representations or hierarchical approaches. These methods could enable the model to handle larger and more complex datasets effectively.

Reviewer hMME4/10 · confidence 4/52024-07-13

Summary

This paper proposes an extension of diffusion models dubbed Equivariant Neural Diffusion, which leverages a learnable forward diffusion process to enhance flexibility. The entire framework has been constructed such that physical symmetry, i.e., equivariance/invariance, of the density is preserved. Experiments are performed on QM9 and DRUGS datasets in the task of molecule generation from scratch as well as controllable generation.

Strengths

1. The presentation is mostly clear and the method is easy to follow. 2. The method has been demonstrated to perform favorably in the controllable generation setting, even when the number of sampling steps is limited to, e.g., 50.

Weaknesses

1. The proposed approach seems to be an incremental combination of geometric diffusion models and neural flow diffusion models. The way of combining these two flavors incurs limited technical novelty. The core design mostly lies in constructing $F\_\varphi$ which is equivariant. 2. The performance on QM9 and DRUGS is a bit marginal compared with the selected baselines. 3. Missing important baselines, e.g., GeoBFN [1], which has shown strong performance on the same task, i.e., molecule generation. Moreover, GeoBFN can also achieve high performance with very few sampling steps, even only 20. [1] Song et al. Unified generative modeling of 3d molecules with bayesian flow networks. In ICLR'24.

Questions

Q1. How does the method perform compared with GeoBFN, especially with different number of sampling steps? Q2. Could the method be applied to systems with larger scales, e.g., proteins?

Rating

4

Confidence

4

Soundness

2

Presentation

2

Contribution

2

Limitations

The authors have discussed the limitations, including scaling, limited application scope, etc.

Reviewer hMME2024-08-12

Thank you for the response. However I do believe the paper would be in better shape by additional efforts in some modifications e.g., adding comparison with advanced baselines (e.g., GeoBFN). I understand that GeoBFN was published ~2months prior to the deadline but the whole experiment would just take around several days and a fair comparison is expected, since you both work on the same dataset and benchmark and claim the advantage of fewer step sampling. Moreover, the work still seems technically incremental. I will slightly increase the score but still hold an opinion that the paper could be further strengthened by either showcasing strong performance against sota methods or exploring novel applications/use cases that constitutes a unique contribution.

Authorsrebuttal2024-08-13

Follow up

We thank the reviewer for engaging in the discussion, and helping us improve the paper further. First, we want to reiterate that a comparison with GeoBFN based on the published results is provided in our rebuttals, and will be included in the final version of the paper — details are provided below for reference. Second, we want to emphasize that END showcases strong performance against SOTA methods, and specifically against GeoBFN. On QM9, END demonstrates a level of performance very similar to that of GeoBFN (see below). On GEOM-Drugs, END clearly outperforms GeoBFN in terms of atom stability (e.g. END’s 50-step atom stability is better than GeoBFN’s 1000-step), while GeoBFN does slightly better in terms of validity. We note that GeoLDM, a very relevant baseline, is shown to lead higher validity than both GeoBFN and END, but that connectivity and geometry quality is subpar compared to END — highlighting that concluding anything from that metric alone is difficult. Finally, as suggested by the reviewer, we will run and collect additional metrics for GeoBFN, and include them in the final version of the paper. However, due to the unavailability of public checkpoints, we can unfortunately not perform that experiment before the discussion period ends. **QM9** | Metrics / Steps | | 50 | 100 | 500 | 1000 | |-----------------|--------|-----------------|----------------|----------------|-----------------| | At. Sta. | END | $98.6 \pm 0.0$ | $98.8 \pm 0.0$ | $98.9 \pm 0.0$ | $98.9 \pm 0.0$ | | | GeoBFN | $98.28\pm 0.1$ | $98.64\pm 0.1$ | $98.78\pm 0.8$ | $99.08\pm 0.06$ | | | | | | | | | Mol. Sta. | END | $84.6 \pm 0.1$ | $87.4 \pm 0.2$ | $88.8 \pm 0.4$ | $89.1 \pm 0.1$ | | | GeoBFN | $85.11 \pm 0.5$ | $87.21\pm 0.3$ | $88.42\pm 0.2$ | $90.87\pm 0.2$ | | | | | | | | | V | END | $92.7\pm 0.1$ | $94.1\pm 0.0$ | $94.8\pm 0.2$ | $94.8\pm 0.1$ | | | GeoBFN | $92.27\pm 0.4$ | $93.03\pm 0.3$ | $93.35\pm 0.2$ | $95.31\pm 0.1$ | | | | | | | | | V x U | END | $91.4\pm 0.1$ | $92.3\pm 0.2$ | $92.8\pm 0.2$ | $92.6\pm 0.2$ | | | GeoBFN | $90.72\pm 0.3$ | $91.53\pm 0.3$ | $91.78\pm 0.2$ | $92.96\pm 0.1$ | **GEOM-Drugs** | Metrics / Steps | | 50 | 100 | 500 | 1000 | |-----------------|--------|----------------|----------------|----------------|----------------| | At. Sta. | END | $87.1 \pm 0.1$ | $87.2 \pm 0.1$ | $87.0 \pm 0.0$ | $87.0 \pm 0.0$ | | | GeoBFN | $75.11$ | $78.89$ | $81.39$ | $85.60$ | | | | | | | | | V | END | $84.6 \pm 0.5$ | $87.0 \pm 0.2$ | $88.8 \pm 0.3$ | $89.2 \pm 0.3$ | | | GeoBFN | $91.66$ | $93.05$ | $93.47$ | $92.08$ |

Reviewer 5dRq2024-08-11

Follow-up Response

I appreciate authors' efforts in addressing my concerns and questions in rebuttal. After reading over authors' rebuttal responses and pdf, I think all my concerns have been addressed so I increased my score. I hope authors will carefully add all rebuttal updates (discussion, analysis and experiment results) to the revised version of paper in the future.

Authorsrebuttal2024-08-12

Follow-up

We want to thank the reviewer for re-evaluating our submission, and providing a positively updated score. We will make sure to include all rebuttal updates in the final version of the manuscript.

Authorsrebuttal2024-08-12

Follow up

Again, we want to thank the reviewer for providing constructive feedback. We believe that we have now addressed the weaknesses raised in the initial review. We are happy to clarify any issue that should remain.

Authorsrebuttal2024-08-13

Follow up ablation on GEOM-Drugs

We want to thank again the reviewer for suggesting us to run additional ablations. As promised in our initial rebuttal, we provide here additional results on the more challenging GEOM-Drugs -- i.e. an ablated model where only the mean is learned (the standard deviation of the conditional marginal is pre-specified and derived from the same noise schedule as EDM). The results are collected in the last row of the table herebelow, and presented with the initial results for better readability. Similarly to the conditional setting, learning only the mean provides a clear improvement compared to the baseline, but is shown to perform slightly worse to the full model across all metrics except validity. | | | At. Stab. [\%] | V [\%] | V$\times$C [\%] | TV$_A$ [$10^{-2}$] | |-----------------|-------|-----------------------------|-----------------------------|-----------------------------|-----------------------------| | Model | Steps | | | | | | EDM | 50 | $84.7_{\scriptstyle \pm.0}$ | $93.6_{\scriptstyle \pm.2}$ | $46.6_{\scriptstyle \pm.3}$ | $10.5_{\scriptstyle \pm.1}$ | | | 100 | $85.2_{\scriptstyle \pm.1}$ | $93.8_{\scriptstyle \pm.3}$ | $56.2_{\scriptstyle \pm.4}$ | $8.0_{\scriptstyle \pm.1}$ | | | 250 | $85.4_{\scriptstyle \pm.0}$ | $94.2_{\scriptstyle \pm.1}$ | $61.4_{\scriptstyle \pm.6}$ | $6.7_{\scriptstyle \pm.1}$ | | | 500 | $85.4_{\scriptstyle \pm.0}$ | $94.3_{\scriptstyle \pm.2}$ | $63.4_{\scriptstyle \pm.1}$ | $6.4_{\scriptstyle \pm.1}$ | | | 1000 | $85.3_{\scriptstyle \pm.1}$ | $94.4_{\scriptstyle \pm.1}$ | $64.2_{\scriptstyle \pm.6}$ | $6.2_{\scriptstyle \pm.0}$ | | END | 50 | $87.1_{\scriptstyle \pm.1}$ | $84.6_{\scriptstyle \pm.5}$ | $68.6_{\scriptstyle \pm.4}$ | $5.9_{\scriptstyle \pm.1}$ | | | 100 | $87.2_{\scriptstyle \pm.1}$ | $87.0_{\scriptstyle \pm.2}$ | $76.7_{\scriptstyle \pm.5}$ | $4.5_{\scriptstyle \pm.1}$ | | | 250 | $87.1_{\scriptstyle \pm.1}$ | $88.5_{\scriptstyle \pm.2}$ | $80.7_{\scriptstyle \pm.6}$ | $3.5_{\scriptstyle \pm.0}$ | | | 500 | $87.0_{\scriptstyle \pm.0}$ | $88.8_{\scriptstyle \pm.3}$ | $81.7_{\scriptstyle \pm.4}$ | $3.3_{\scriptstyle \pm.0}$ | | | 1000 | $87.0_{\scriptstyle \pm.0}$ | $89.2_{\scriptstyle \pm.3}$ | $82.5_{\scriptstyle \pm.3}$ | $3.0_{\scriptstyle \pm.0}$ | | **END (mean only)** | 50 | $85.6_{\scriptstyle \pm.1}$ | $87.8_{\scriptstyle \pm.2}$ | $66.0_{\scriptstyle \pm.4}$ | $7.9_{\scriptstyle \pm.0}$ | | | 100 | $85.8_{\scriptstyle \pm.1}$ | $89.9_{\scriptstyle \pm.1}$ | $73.7_{\scriptstyle \pm.4}$ | $6.1_{\scriptstyle \pm.1}$ | | | 250 | $85.7_{\scriptstyle \pm.1}$ | $91.2_{\scriptstyle \pm.2}$ | $77.4_{\scriptstyle \pm.4}$ | $5.0_{\scriptstyle \pm.1}$ | | | 500 | $85.8_{\scriptstyle \pm.1}$ | $91.6_{\scriptstyle \pm.1}$ | $78.6_{\scriptstyle \pm.3}$ | $4.8_{\scriptstyle \pm.1}$ | | | 1000 | $85.8_{\scriptstyle \pm.1}$ | $91.8_{\scriptstyle \pm.1}$ | $79.4_{\scriptstyle \pm.4}$ | $4.6_{\scriptstyle \pm.0}$|

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC