Author Response - Part 2
**5. Benchmark with baselines with the same amount of training time**
We first clarify that in Table 1, we follow a standard criterion in the literature, i.e., ensuring the same number of fine-tuning tokens across both the baselines and our method.
Following your suggestion, we further train the baselines with the same amount of training time as our method, which corresponds to 20k training iterations on Alpaca-gpt4. As shown in the table below, we report the MMLU accuracy as well as the average accuracy across 7 tasks, following the task list of Table 1. We observe that (1) the baselines fine-tuned with 20k iterations generally maintain comparable performance with those trained with 10k iterations, potentially because 10k iterations is sufficient for fine-tuning on Alpaca-gpt4; and (2) our method still outperforms all baselines across various tasks.
| **Remaining Ratio** | **Method** | **Training Iterations** | **Training Time** | **MMLU (%)** | **Average Acc (%)** |
|:---:|:---:|:---:|:---:|:---:|:---:|
| 80% | FLAP | 10k | 15h | 40.21 | 60.98 |
| | | 20k | 30h | 38.62 | 61.40 |
| | Shortened LLaMA | 10k | 15h | 26.45 | 58.72 |
| | | 20k | 30h | 26.55 | 59.63 |
| | Ours | 10k | 30h | **40.70** | **62.29** |
| 65% | FLAP | 10k | 15h | 33.28 | 56.12 |
| | | 20k | 30h | 35.4 | 55.93 |
| | Shortened LLaMA | 10k | 15h | 24.89 | 52.57 |
| | | 20k | 30h | 24.70 | 53.95 |
| | Ours | 10k | 30h | **36.00** | **56.96** |
| 50% | FLAP | 10k | 15h | 27.67 | 51.12 |
| | | 20k | 30h | 27.55 | 51.32 |
| | Shortened LLaMA | 10k | 15h | 24.76 | 47.35 |
| | | 20k | 30h | 25.10 | 49.22 |
| | Ours | 10k | 30h | **30.60** | **52.19** |
**6. The number of remained depth and width in Table 1**
Thank you for the suggestion! The (depth, width scale) for AmoebaLLM with 80%/65%/50% remaining ratios in Table 1 are (30, 0.875)/(28, 0.75)/(22, 0.75), respectively. We will follow your suggestion to add this information to Table 1 to improve readability.
---
**7. The detailed subnet selection strategy**
The profiling results in Section 3 are purely for motivating the problem. Here is our current subnet selection strategy: we adopt a hierarchical search strategy to deliver subnets from our design space that satisfy the target efficiency constraint, e.g., 50% weight remaining ratios in Table 1 of our manuscript, while maximizing the achievable accuracy. Specifically, we first perform a coarse grid search across uniformly spaced depth and width settings based on a small calibration set, e.g., 20 samples from the MMLU dataset, to identify the subnets that satisfy the given efficiency constraint with maximized accuracy. Next, we perform a more fine-grained grid search within depth/width ranges surrounding the optimal subnet identified in the coarse grid search stage. This process typically evaluates 40 promising subnets and takes no more than 10 minutes on an NVIDIA A5000 GPU.
We empirically find that the above strategy works well at the scale of our target problem and can already deliver high-quality subnets that outperform previous compression methods. We also note that more complex subnet selection strategies, such as the evolutionary search adopted by [11]-[15] cited in our manuscript, can also be employed, which will be our future work.
We will clarify this in the final version.
---
**8. The rank of LoRA**
We follow QLoRA [1] and adopt a rank of 64. We will add this information to the final version.
[1] “QLoRA: Efficient Finetuning of Quantized LLMs”, T. Dettmers et al., NeurIPS’23.