AI for Interpretable Chemistry: Predicting Radical Mechanistic Pathways via Contrastive Learning

Deep learning-based reaction predictors have undergone significant architectural evolution. However, their reliance on reactions from the US Patent Office results in a lack of interpretable predictions and limited generalization capability to other chemistry domains, such as radical and atmospheric chemistry. To address these challenges, we introduce a new reaction predictor system, RMechRP, that leverages contrastive learning in conjunction with mechanistic pathways, the most interpretable representation of chemical reactions. Specifically designed for radical reactions, RMechRP provides different levels of interpretation of chemical reactions. We develop and train multiple deep-learning models using RMechDB, a public database of radical reactions, to establish the first benchmark for predicting radical reactions. Our results demonstrate the effectiveness of RMechRP in providing accurate and interpretable predictions of radical reactions, and its potential for various applications in atmospheric chemistry.

Paper

References (52)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer G4db5/10 · confidence 4/52023-06-27

Summary

The paper proposes a set of methods for modeling chemical reactions that involve radicals during the reaction process. The authors first introduce the current landscape of chemical reaction datasets based primarily on the USPTO and discusses the USPTO shortcomings in terms of reaction interpretability and its inability to showcase reaction involving multiple steps. Next the authors briefly describe the RMechDB dataset, which does contain pathways with radicals, followed by a description of their methods. The methods are based on OrbChain, which provides a standardized way of describing chemical reactions with radical and arrow pathways. After that, the authors introduce their predictive methods, including two-step prediction, plausibility ranking, contrastive learning with a reaction hypergraph and text-based sequence to sequence models. The authors conduct experiments for all the aforementioned methods to further understand their capability in accurately describing radical reaction pathways on the RMechDB dataset. The performances of the different methods vary across different settings of the conducted experiments. The authors then provide a pathway search example, further description of their package and a conclusion.

Strengths

The paper provides has the following strengths: * Originality: The papers provides a new perspective on chemical reaction modeling that involves radicals and is also more interpretable for classically trained chemists. * Quality: The paper describes and analyzes four different and relevant methods for the radical modeling problem and provides clear motivations for their importance. * Clarity: The paper motivates the problem they address quite and describe the necessary background. * Significance: Expanding the capabilities of machine learning models to provide more interpretable reaction models with more steps in the reaction process could have significant impact on various chemistry related problems.

Weaknesses

The paper could be further improved: * Providing a clearer description of the context of the results related to original problem the authors motivated. How well do the described methods provide more interpretability to chemical reaction modeling? How do the metrics the authors measure relate to that original premise? [quality, clarity, siginificance] * The authors only briefly describe pathway search, but provide little context for what their results mean. What does a recovery rate of 60% imply? How does the reaction tree look like and how interpretable is it? [clarity] * The authors refer the reader to the appendix very often, which I think contains a lot of significant information needed to fully understand the experimental results. I recommend putting more of that information in the main paper. [quality, clarity significance] * The authors only provide a brief description of the RMechDB dataset and its unclear if that paper had any modeling methods the authors could compare their proposed methods to. Further clarification on this would be helpful. [clarity]

Questions

* Could you clarify if the RMechDB paper provided any modeling methods? * Would it be possible to provide more context for the results, ideally in the figure and tables themselves, to further understand the experiments? How good does a Top1, Top2, Top5 score need to be practically useful, for example? * Is it possible to have text-based methods also express the intermediate steps in reactions involving radicals? Why or why not?

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

2 fair

Contribution

3 good

Limitations

The authors do not provide a detailed discussion on limitations. A discussion on limitations would make the paper stronger.

Reviewer PKyZ6/10 · confidence 4/52023-07-07

Summary

The authors provide a reaction predictor system that provides an accurate and interpretable prediction of radication reactions. Due to the lack of training data, there is a dearth of reaction predictors for radication reactions. The authors present 3 deep-learning-based approaches. The first approach is a two-step process that identifies possible reactive sites and then ranks the reactive site pairs. The second approach uses a contrastive learning approach to identify the most reactive site pairs. Finally, the authors also show a transformer-based approach to perform sequence-to-sequence translation from products to reactants. In the two-step, OrbChain approach, the authors present a GNN-based approach to identify reactive sites and a siamese network-based approach to rank the plausible reactive sites. Multiple reaction representations to perform plausibility ranking. The contrastive learning approach also uses a GNN and both a custom atom pair representation and a hypergraph representation are evaluated. Finally, a pre-trained MolGPT on USPTO dataset is used as well. The authors find that the graph-based methods outperform MolGPT and the contrastive learning methods yield the most accurate results.

Strengths

- The authors compare multiple models to show the efficacy of different types of models such as GNNs and text-based Transformers for reaction prediction - The authors also use multiple representations and model architectures for a very thorough evaluation of the proposed reaction

Weaknesses

- It is not clear how or which of the three algorithms described is used in RMechRP. - The presentation of the paper could be improved. There are 3 approaches described with multiple models and representations for some approaches. A short summary of the findings and comparisons or a visualization of the approaches could significantly improve the presentation

Questions

- Line 100: How is the arrow-pushing mechanism A represented in OrbChain? - In Table 2, what is the Atom Fingerprint method? Morgan fingerprint? ECFP? - What is the loss function for the contrastive learning approach? (Might be in the appendix) - Link 334: contrstive -> contrastive?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

2 fair

Contribution

3 good

Limitations

- Limited comparison as the authors don’t include a related works section so it is difficult to contextualize the scope of their current work to the field. - Is there a reason the MolGPT model could not be trained on the new RMechDB dataset for evaluation rather than only fine-tuning?

Reviewer URj37/10 · confidence 5/52023-07-20

Summary

- a new model is described for prediction of radical chemical reactions - the model is trained on a dedicated database of radical reactions for atmospheric chemistry, an important application - several, reasonable baselines are evaluated

Strengths

- reasonable, state of the art ML modelling (contrastive learning, attention GNNs, reasonable reaction representations inspired from molecular orbitals, building on previous work by Baldi's group) - reasonable strong baselines (transformers) - compelling results - important application

Weaknesses

- other baselines, like MEGAN https://pubs.acs.org/doi/abs/10.1021/acs.jcim.1c00537 or https://www.nature.com/articles/s42256-022-00526-z could be considered ### Related work Several references in the introduction are not correct: The Cao & Kipf MolGAN paper should be removed, because it does not deal with chemical reactions. similarly, the Rogers et al ECFP does not deal with reaction prediction, and should be removed in the intro. On the other hand, the Segler et al paper should be cited as an ML paper. The ELECTRO paper by Bradshaw et al should be added. https://arxiv.org/abs/1805.10970 contrastive learning to distinguish between plausible and implausible reactions has already been used in https://www.nature.com/articles/nature25978 (called in-scope filter there), which should be referenced as well

Questions

no questions, this is a solid, straightforward paper in my opinion

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

4 excellent

Presentation

3 good

Contribution

4 excellent

Limitations

n/a

Reviewer n7xM6/10 · confidence 4/52023-07-21

Summary

Authors present two models that predicts radical chemistry reactions. The first model 'OrbChain' is comprised of two components 1) one GNN model for predicting pairs of reacting atoms/groups, 2) a model which ranks the plausibility of these pairs. The second model is a a fine-tuned Rxn-Hypergraph model, adapted for the task of predicting radical mechanism by using an atom classifier model.

Strengths

+ Selects an interesting problem domain, specifically radical based chemistry. + Authors plan to open source the Radical Mechanistic Prediction model and release software for easier use. + Compares against relevant baselines, such as fingerprint representations, MolecularTransformer.

Weaknesses

The presentation of results could be more clear. In particular: - It would be helpful to have a clear statement of the key contributions provided by this work. The list of desiderata provided at the end of section 2 are important, but I believe that these properties have already been provided by previous models, especially references [20, 21] for radical reactions. - If I understand correctly, OrbChain is the name of the two part model, but components of the model are still used for the second modeling approach using the fine tuned Rxn-Hypergraph model. - The table formatting makes the results somewhat difficult to parse, it would be helpful to have more spacing between the caption and the table, and for the - Table 3, where there is one column with 'AP \n Morgan2 \n TT' and it wasn't immediately obvious that these are different molecular descriptors. - For Figure 3, I believe the reaction type should be 'Homolysis' rather than homolyze. - For Figure 3, It would be helpful to have a sense of the number of reactions in each class to better compare the relative performance by the model between reaction classes. - There are several typos in the manuscript, e.g. a missing close parenthesis in lines 48-49 of page 2, 'weather' instead of 'whether' on line 188 on page 5, some tense mismatches. Please review for grammar errors. In Section 2, authors state that 'None of the currrent reaction predictors can offer ... chemical interpretability, pathway interpretability, or balanced atom mapping'. There are actually several models that provide interpretability for reaction mechanisms/reaction type. In addition to the works on radical mechanism prediction cited by the authors as references 20 and 21: - In https://arxiv.org/pdf/1805.10970.pdf, Bradshaw et al. predict electron pair pushing mechansims with a generative model. - The MolecularTransformer model has also been shown to provide atom mapping by visualizing attention weights (https://arxiv.org/pdf/2012.06051.pdf, Figure 2)./ For text based models such as Molecular Transformer, could you quantify the percentage of reaction predictions that suffer from a 'balance problem'? It isn't clear to me that this is a big issue with MolecularTransformer or other text based models.

Questions

- Could you provide more information on the train test split used in the RMechDB? Are there any splits that are used to measure generalizability of predictions to out of domain reactions (e.g. by reaction type, atom types, structure similarity) - Since interpretability is one of the benefits highlighted by this modeling approach, it would be interesting to see more examples of mechanism prediction by the proposed model in the main text. - For the Pathway Search task, the results say that 60% of the reactants were found in the expanded reaction trees (of 10 step mechanisms). How well do current non-ML techniques perform on this task? How many of the proposed reactants are false positives; are false positive predictions of reactants detrimental to the problem prediction? - What is the final model used in the RMechRP Software?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

2 fair

Contribution

3 good

Limitations

I do not identify any negative societal implications. By my understanding, the presented model is intended to be limited for only radical based reactions, as it is trained on this domain of reactions, and not for other types of chemical reactions.

Reviewer n7xM2023-08-11

Thank you authors for their explanations and efforts to make the work more accessible. The addition of Figure 1 in the attachment is very helpful for laying out the problem and the proposed solutions. I think that the authors could do more to expand either Figure 1 or the caption to Figure 1 to better highlight the contributions by this work. Based on my understanding of your responses and rereading the paper, my understanding of the main contributions of this work are: - Benchmarking several methods on the radical specific dataset RMechDB (ref. 25). The transformer model, RxnHypergraph are previously published models, but authors here introduce new features to make use of the RxnHypergraph model. Of the proposed three methods, only the two-step prediction method provides both chemical interpretability (i.e. an assignment of molecular orbitals), but all methods contribute pathway interpretability. - OrbChain is a new GNN model presented in this work that can 1) Classify reaction types, 2) Predict reaction outcomes using the reaction types. The method here is somewhat similar to the reaction prediction algorithms published ref. 20 (Learning to Predict Reactions). I do not follow the authors' point in the rebuttal about how the model in this reference does not provide chemical interpretability -- If I understand this work correctly, I believe specific one-step mechanisms are predicted by the model in the form of identifying electron 'sources' and electron 'sinks' and solving a matching problem. If I understand correctly, the main difference with OrbChain is that it is a GNN model, and covers a much wider scope of reactions because of the use of RMechDB as training; is this correct? With the changes made by the authors to help clarify the contributions of this work, and the changes proposed by the authors to improve readability/typos, I have edited scores given in my original review. Question specifically about Figure 1 in the attachment: - What is meant by OrbChain generating 'Labels'? Does this mean reaction classification labels?

Authorsrebuttal2023-08-15

We thank the reviewer for their constructive comments. We also appreciate their willingness to adjust their score. Using their second set of comments we have further revised our manuscript. **Regarding Figure 1 in the attachment** * We agree with the reviewer that the caption is perhaps too concise. As a result, in the revised version, we have expanded the caption to read: “This is a schematic depiction of the prediction problem, the processing tool (OrbChain), and the three approaches. The three approaches are: Two Step Prediction, Contrastive Learning, and Text-Based. The first two approaches use OrbChain to find the reactive orbitals for training and to form the products during inference.” **Regarding the first question on OrbChain** * We wish to clarify that OrbChain is not really a GNN, but rather a reaction processing tool that is used by the GNN models (see Equation 1 describing OrbChain). This processing step assigns labels to molecular orbitals and their atoms before training the GNNs. This also answers the second question about OrbChain raised by the reviewer: OrbChain provides labels at the level of orbitals and atoms, not at the level of reactions. The text-based models do not need OrbChain because they operate directly on the text representations, not the orbitals. This reviewer is correct in stating that the use of the RMechDB dataset significantly expands the scope of the reactions.

Reviewer G4db2023-08-14

Thank for the additional details

The authors have clarified many of my major concerns and I have adjusted my score accordingly.

Authorsrebuttal2023-08-17

We appreciate the reviewer's constructive comments and their willingness to adjust their score. We welcome further input to enhance our paper's quality in the time ahead.

Authorsrebuttal2023-08-19

As the deadline is approaching, we are keen to ensure that our responses addressed all the concerns raised in your review. Based on your valuable feedback, we have made substantial revisions to improve the quality and clarity of our submission. Therefore, could you kindly consider updating your review or score to reflect these improvements? If there are still any unresolved concerns or areas that need further clarification, please let us know so that we can address them promptly.

Reviewer PKyZ2023-08-21

Thank you for the clarification. I believe that I will stay with my current rating of 6.

Authorsrebuttal2023-08-21

We appreciate the reviewer's constructive comments. We believe addressing these comments has led to significant improvements in both quality and clarity of our submission.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC