We sincerely thank you for your insightful feedback. Below are our detailed responses to your concerns:
**1. More domain-specific baselines**
Based on your recommendations, we have expanded our set of baselines to include Text+Chem T5[1], Galactica[2], and MolT5[3] as these models have undergone instruction tuning, aligning them more closely with the objectives of our study.
Regarding the baselines [4,5,6], we have chosen not to include them in our comparison as they focus on pretraining purely on molecular data without any instruction-following. This difference in approach makes a direct comparison somewhat unfair. However, we would like to clarify that our evaluation methods for reagent prediction, forward reaction prediction, and retrosynthesis are consistent with those used by Text+Chem T5[1], which we hope alleviates some of your concerns.
For protein-related tasks, the most closely related baseline we found was [7]. However, due to the unavailability of their code, we were unable to include it in our direct comparison. Instead, we have incorporated Galactica[2] as a baseline for these tasks, considering its relevance and accessibility.
The updated results reflecting these expansions can be found in **Tables 3 and 4**, as well as **Figures 5, 6, 8, 9, 10, and 11** in our revised manuscript.
**2. Lack of data processing details**
Due to page constraints, a detailed description of our data transformation process, from the original sources to instruction data, was not feasible within the main text. However, to ensure transparency and reproducibility, we have included an extensive explanation in **Appendix B (Pages 15-22)**. This section thoroughly details each task's definition, data sources, and the methodology employed to convert this data into the intended instruction format. We believe this provides a clear and comprehensive understanding of our data processing approach. Moreover, we have also added further details on quality control measures **(Highlighted in Section 3.2, Page 4)**.
**3. Insufficient tasks**
We apologize for any misunderstanding and would like to clarify that our experiments comprehensively cover the 17 tasks listed in Figure 3. These include:
- **Molecule-oriented Tasks:** Molecular Description Generation (Figure 5), Description-guided Molecule Design (Table 4), Forward Reaction Prediction (Table 4), Retrosynthesis (Table 4), Reagent Prediction (Table 4), Property Prediction (Table 3).
- **Protein-oriented Tasks:** Domain/Motif Prediction (Figure 5), Functional Description Generation (Figure 5), Protein Function Prediction (Figure 5), Catalytic Activity Prediction (Figure 5), Protein Design (Figure 9)
- **Biomolecular Text Tasks:** Open Question (Figure 6), Multi-choice Question (Figure 6), Chemical-protein Interaction Extraction (Figure 6), Chemical-disease Interaction Extraction (Figure 6), Chemical Entity Recognition (Figure 6), True or False Question (Figure 6).
We believe that the variety of tasks covered by Mol-Instructions provides a substantial representation of common biomolecular challenges. Nevertheless, we recognize that there are aspects of biomolecular research that we have not yet explored. These represent opportunities for further enhancement in future research, and we are committed to continually expanding and refining our dataset to encompass an even wider range of biomolecular challenges.
Thank you again for your valuable suggestions! Your input has significantly contributed to enhancing the fairness of our experimental comparisons and the overall completeness of our paper.
[1] Unifying Molecular and Textual Representations via Multi-task Language Modelling. In ICML 2023.
[2] Galactica: A Large Language Model for Science. 2022.
[3] Translation between Molecules and Natural Language. In EMNLP 2022.
[4] Root-aligned SMILES: a tight representation for chemical reaction prediction. In Chemical Science 2022.
[5] Chemformer: a pre-trained transformer for computational chemistry. In Mach. Learn.: Sci. Technol. 2022.
[6] Reagent prediction with a molecular transformer improves reaction data quality. In chemical science.
[7] ProteinDT: A Text-guided Protein Design Framework. 2023.