Second-round Responses (I)
Thanks for acknowledging our theoretical and empirical **Response to Weakness 1 and Limitation 1: Overall invertibility of entire neural networks is well-supported**.
For the new questions, we respond with the following answers (**NQ** stands for new question):
## NQ.1. BPTT + MIM4DD is lower than BPTT?
Thank you for highlighting this discrepancy. It's essential to note that the results for BPTT we have reported are based on our strict reproduction using BPTT's official codebase. Despite our meticulous adherence to the methodology, we were unable to replicate the exact results claimed by BPTT. Moreover, based on our examination of all 17 papers citing BPTT, none of them refers to BPTT's reported results, which further suggests that other researchers might also be facing challenges in reproducing those numbers. Thus, in our paper, we decided to use the results we obtained from our reproduction since they still remain competitive and within the state-of-the-art range. We appreciate your understanding and will clarify this in our revision.
Additionally, in light of Reviewer pJjQ's comments, we've made efforts to further compare our method using the DREAM (ICCV2023) framework, resulting in the achievement of new state-of-the-art results. DREAM is a stronger baseline than BPTT. **More details can be found in our response to Reviewer pJjQ.**
### ( Using DREAM (ICCV2023) [R.1] as a baseline framework, we reached a new SOTA!)
**Why DREAM was chosen for additional experimentation:**
- Despite not being a required comparison according to NeurIPS policy, DREAM represents the cutting-edge in dataset distillation, and we aim to remain at the forefront of this research area.
- Given the limited rebuttal timeframe, DREAM's efficiency and clear codebase provided an ideal setting for our experiments.
- DREAM's codebase is compatible with gradient and feature matching-based dataset distillation frameworks. Incorporating these provides a comprehensive response to concerns raised about these aspects in our previous evaluations.
Building upon DREAM's framework, we integrated our MIM4DD module. The cluster-wise DD approach of DREAM was retained, with our contrastive aligning module enhancing DREAM’s match loss component. All experimental settings strictly adhered to DREAM's parameters, ensuring that only our module contributed to any observed variations.
**Experimental Results:**
Adding our method MIM4DD on DREAM.
| Method | CIFAR10 IPC-1 | CIFAR10 IPC-10 | CIFAR10 IPC-50 | CIFAR100 IPC-1 | CIFAR100 IPC-10 |
|-------|-------|-------|-------|-------|-------|
| DREAM [R.1] | 51.1±0.3 | 69.4±0.4 | 74.8±0.1 | 29.5±0.3 |46.8±0.7 |
| DREAM + MIM4DD | 51.9±0.3| 70.8±0.1 | 74.7±0.2 | 31.1±0.4 | 47.4±0.3 |
Top-1 accuracy of test models trained on distilled synthetic images on **TinyImageNet**.
| IPC | Ratio % | DM [39] | MTT [5] | DREAM [R.1] | DREAM +MIM4DD | Whole |
|-------|-------|-------|-------|-------|-------|-------|
| 1 | 0.017 | 3.9±0.2 | 8.8±0.3 | 10.0±0.4 | 11.2±0.2 | 37.6±0.4 |
| 10 | 0.17 | 12.9±0.4 | 23.2±0.2 | 23.9±0.4 | 24.8±0.3 | 37.6±0.4 |
These results underline that MIM4DD, when integrated to DREAM, further enhances performance. While hyper-parameters weren't exhaustively fine-tuned, the results reflect MIM4DD's versatility across different dataset distillation frameworks.
In conclusion, the enhancement of DREAM's results with our MIM4DD module attests to its efficacy and adaptability. We appreciate the reviewer's feedback, which provided an avenue for us to further highlight the method's robustness and relevance in contemporary DD research.
## NQ.2. Results on TinyImageNet
Please refer to NQ.1. We use a new SOTA codebase to realize the experiments on TinyImageNet.
**reference**
[R.1] DREAM: Efficient Dataset Distillation by Representative Matching, ICCV 2023