Response to comment by Reviewer Lxfs
We thank the reviewer for their response and are happy to provide further clarification of the concerns mentioned.
>**The validation of scaling laws in CV and NLP domains largely relies on unlabeled data and self-supervised tasks [1,2,3] [...]. [A] more relevant approach for studying scaling is under an unlabeled setting [e.g.,] [4].**
Unsupervised approaches are an interesting avenue of research in certain areas of the molecular domain. As suggested, Uni-Mol models [4,5] show promising results for modeling molecules in 3D space and for finetuning to physics-based downstream tasks like QM9.
We would like to remind the reviewer that we follow a significantly different approach based on modeling 2D graphs with different applications & downstream tasks, where unsupervised models have not yielded comparable results.
In the context of CV/NLP [1,2,3] we also point out the lack of equally label-rich datasets that explains why supervised pretraining is not a focal point in those domains.
With respect to Uni-Mol2 [4], we note that it has been only finetuned to physics-based tasks (QM9, e.g., homo-lumo gap), that highly depend on 3D coordinates and graph structure. [4] is specialized on understanding graph structure in 3D space, which leads the approach to perform well on such tasks.
>We hope that our work paves the way for an era where foundational GNNs drive pharmaceutical drug discovery. [abstract of our submission]
Our objectives are different: predicting properties that drive drug discovery like ADMET and binding predictions that strongly rely on understanding the biochemical space beyond graph structure. Uni-Mol2 provides no results on any shared or similar downstream task.
While the initial Uni-Mol model [5] provided some results for ADMET downstream tasks, the derivation of the results lacks transparency and is not reproducible (despite our best efforts).
Uni-Mol's finetuning compares to GraphMVP [6] (published one year before Uni-Mol), which we also compare to in our work, closely following their experimental setup. Uni-Mol outperforms GraphMVP but reports performance directly taken from [6], despite clearly using a different dataset splitting technique in their work. We reiterate that data splits have a significant impact on empirical results in the context of molecular scaffold splits, especially in the low-sample regime of the MoleculeNet. Further, [5] does not provide sufficient information for reproducibility. We have rigorously analyzed the papers and code of both works and provide details in a separate comment below.
Overall, the learnings from our scaling study (width, depth, #molecules, #labels per molecule, composition of pretraining data mixture) lead to the derivation of a foundational GNN that has set a new standard in the various competitive downstream task benchmarks considered here. We kindly request the reviewer to provide more context on their assessment of our work as “a trivial contribution” in light of the above discussion.
We would be happy to clarify any further questions the reviewer may have.
>**In the biological domain, obtaining labeled data is more expensive and challenging compared to the CV or NLP domains. [...] The effectiveness of this approach may only depend on the relevance of supervised tasks between pre-training and fine-tuning.**
We respectfully disagree with this assessment. Adequate labeled data is readily and publicly available at large scale with our pretraining only using a small fraction of the available data.
We recall that PCBA_1328 is only a small subset of the PubChem database (only considering datasets with binary tasks with at least 6k molecules) and larger alternatives for PCQM4M exist (e.g., PM6 dataset in [7] that is more than 20x bigger).
We also recall that our results show no signs of a data bottleneck. Instead, our molecule scaling suggests, downstream task performance does not deteriorate much when pretraining on a smaller fraction of molecules, e.g., 25% or 50% (Fig. 2 of submission). This relates back to the importance of the number of labels (i.e., label scaling in submission) per molecule already discussed in our initial response.
The objective of our work, as in CV/NLP, is to obtain informative embeddings for domain-specific downstream applications. For drug discovery, our work suggests supervised pretraining is a suitable avenue, thanks to the availability of large labeled data sources. As there is no overlap between the pretraining and finetuning tasks, our downstream evaluation solely evaluates the information content of our learned representations (similar to other domains).
We are happy to provide more information if further questions arise.
[1,2,3,4] as referenced in previous comment of Reviewer Lxfs
[5] Uni-Mol
[6] Liu et al, Pre-training Molecular Graph Representation with 3D Geometry, ICLR 2022
[7] Beaini et al, Towards Foundational Models for Molecular Learning on Large-Scale Multi-Task Datasets, ICLR 2024