Summary
Thank you for your effort to provide comprehensive benchmarks for clinical practice. One of the strongest points is that the benchmark was designed based on several departments and diseases. Another point is the ClinicalAgent can perform an end-to-end for real-world clinical diagnostic practices.
However, the reviewer found major concerns in the manuscript. Those concerns concern the proof that the study's claims are not convinced, the reproducibility of the proposed benchmark, and the datasets.
Weaknesses
proof that the study's claims are not sufficient,
the reproducibility of the proposed benchmark,
and the datasets.
Questions
1. It is mentioned that “LLMs still struggle to meet the strict requirements for accuracy and reliability in the medical field and face many challenges in clinical applications” ===> However, the current work does not any proof to prove that it can deal with this challenge?
Please provide specific evidence or comparative results demonstrating how the approach improves accuracy and reliability compared to existing benchmarks or LLMs in medical applications.
2. It is mentioned that “existing medical evaluation benchmarks face the risk of data leakage or contamination. And “We ensure that ClinicalBench does not have data leakage” ===> The reader can not find the detailed information to convince the claims from the author.
Please describe your methods to prevent data leakage, such as data collection procedures, preprocessing steps, or validation techniques used.
3. Existing evaluation methods are limited to multiple choice questions, which do not align with the real-world diagnostic scenarios ===> Propose Generative QA. The reviewer is not sure of the reliability of the questions; how are the composed validated sets validated by the healthcare professional practitioner? How can we follow this strategy and prepare for our in-hospital domain dataset?
4. The data processing, is not sufficient enough ===> How to deal with the preprocessing to have a complete structured data for the training and inference? For example, how to deal with numeric attributes from inside the notes?
5. Lacking of the demographic statistical analysis of the data (sex, age, etc.) Therefore, we are not certain of the generalization of the benchmark for the generated output from the agent? Especially, with the number of 1500 samples, it is actually not a large enough dataset, so we need to have more detailed how the dataset covers patient demographic statistic?
Please include a detailed demographic breakdown of their dataset, including age ranges, gender distribution, and other relevant factors. Additionally, please discuss how you ensured adequate representation across different demographic groups given the relatively small sample size.
6. We should expect to have the comparative analysis between the proposed benchmark, and the existing bechmarks from the literature? If not, it is not convinced to confirm the effectiveness of the proposed approach?
7. The most critical is the experiment setup in detailed so that the reproducibility can be made? Hyperparamters, fine-tuning approaches, etc….?
Please provide a detailed appendix or supplementary material containing complete hyperparameter settings, specific fine-tuning procedures, data preprocessing techniques, evaluation metric implementations, and code or pseudocode for key algorithms.
8. The experimental evaluation is completed through API calls and 8 NVIDIA A6000 GPUs ====> Should we have a table that compare the computational resource for the training between the proposed benchmark with different LLM models versus with existing benchmarks? Based on that, the interested readers should expect to estimated how much computational resource (training times, inference times, FLOPs, etc..) before they want to replicate the benchmarks and selected the LLM models in advance.
Please Include a comparative table showing computational resources (e.g., training times, inference times, FLOPs) for their benchmark across different LLM models and for existing benchmarks. This would help readers estimate resource requirements for replication or model selection.
Ethics concerns
In case the study is accepted for publication, how can we make sure that the datasets was obtained, prepared.
Especially, the composed set of generative QA from the study was not clearly explained or validated. Therefore, it is impossible for the interested reader want to follow the strategy to design them for performing the benchmark.