Reply to rebuttal comments 1/2
> Since CARL does not have the results on ImageNet, is the parameter of CARL tuned for ImageNet.
As the reviewer mentioned, CARL was not trained on ImageNet. We did the hyperparameter search for the number of prototypes due to our experience executing CARL. When CARL was trained with 10000 prototypes, training collapsed. We kept the entropy weight as 2 (following CARL’s guidelines for other datasets) and reduced the number of prototypes to 3000 (where we observed no collapse).
To fully address the reviewer’s concerns, we pre-trained multiple setups of CARL on ImageNet for 100 epochs. We fixed the number of prototypes as 10000 and ablated different values of lambda (entropy weight). Results are below.
\begin{array} {rrrrrr}
\hline \lambda &1&2&3&4&5\\\\
\hline
CARL&C&C&65.1&64.8&63.9\\\\
\hline
\end{array}
As discussed, using 10000 prototypes causes collapse (**C**), and the workaround is to increase the entropy weight. For reference, CARP 100 epoch model with 2 views achieves 69.7%.
As discussed in Section 3.1, learning that many prototypes at once with the entropy term, as in CARL, leads to underperformance.
> Table 1 [...] CARL shows that it works well with 10,000 prototypes on CIFAR-100. Therefore, 10,000 prototypes should not be a big issue on ImageNet.
We respectfully disagree with the reviewer on this point. The training of CARL on a much simpler dataset with only 100 classes cannot be extrapolated to a more complex dataset with 1000 classes (10 times more). Moreover, CARL’s paper discusses the relation of the number of prototypes and the number of classes (Fig. 3 in CARL’s paper) which is in line with the observed behavior in our experiments. **In CARP, we show that there is a lack of generalization in the usage of prototypes and propose a solution.**
> [...] it is better to have CARP on CIFAR-100 for a fair comparison.
The results comparing CARP and CARL on CIFAR-100 are below.
\begin{array}{rrrrrrr}
\hline
& Cifar10&&Cifar100&&STL10&\\\\
\hline
Ep. &100&200&100&200&100&200\\\\
\hline
CARL&73.39&78.94&42.91&48.85&76.9&81.95\\\\
CARP&74.84&79.52&44.67&50.64&78.05&82.44\\\\
\hline
\end{array}
We pre-trained CARP on CIFAR-10/100 and STL-10 following the CARL’s guidelines (Table 2 on CARL’s paper). We report average top-1 (linear probing) results across 3 independent runs (same as CARL). The number of prototypes was set to 100, 300, and 300 for Cifar-10,-100, and STL-10, respectively (same as CARL), and the partition size was set to 50 for all datasets (no tuning was done to select this partition size). We report models trained for 100 and 200 epochs. CARP outperforms CARL in all datasets. These experiments showed that CARP works well for simpler datasets as well as in ImageNet. By default, CARP uses $\lambda=1$ while CARL uses larger values to avoid collapse, e.g., $\lambda$=2, which explains CARP’s improvements.
> Moreover, the proposed method eliminates the parameter for the entropy while introducing an additional parameter for the partition size, which may not save tuning efforts.
We agree with the reviewer that our method introduces a new hyperparameter (the partition size) and avoids the tuning of the entropy weight. However, **the entropy weight and the partition size have different tuning difficulties**. As discussed in Section 3.1, without the random partition pretext task, training collapses if the entropy term is too small, and accuracy is suboptimal if the entropy term is too high. On the other hand, CARP is robust to the choice of partition size, as shown in Table B.2 (appendix).
> [...] Compared with CARP with multi-crop augmentation as in Table 1 in the submission, the gap [with CoKe] is quite marginal, i.e., 0.3% on Avg@k. A fair comparison is essential to draw any conclusion.
As pointed out by the reviewer, when comparing CoKe with CARP with multi-crop, CARP still outperforms CoKe by a small margin (even though CoKe was trained for 800 epochs and CARP for 400). To clear the reviewer's concerns, we extended the previous comparison against CoKe to include all instances of CARP and an additional instance of CoKe trained for 800 epochs w/o multi-crop. CARP consistently outperforms CoKe by large margins.
\begin{array}{lrcrrrrrrrrrrr}
\hline
Method&Ep&Pets&Flowers&Aircraft&Cars&Country&Food&STL>SRB&Avg @k&&&\\\\
&&&&&&&&&&10&20&100&200\\\\
\hline
CoKe &1000&81.3&75.3&29.3&22.6&13.3&60&95.7&64.2&55&55.2&54.1&53.4 \\\\
CoKe (mc) &800&79.5&79.5&27&22.4&14.6&58.9&95.7&60.4&57.1&57.4&57.1&56.3\\\\
CARP&200&86.8&78.2&38.9&29.8&12.2&58.4&95.5&73.7&59.2&59.2&58.5&57.9\\\\
&400&86.8&80&\textbf{42.1}&\textbf{33.5}&12.3&58.4&95.9&75.3&60.4&60.5&59.7&59.2\\\\
&800&\textbf{87.3}&\textbf{81.2}&41.1&33.2&13.6&61.2&\textbf{97}&\textbf{76.4}&\textbf{61.2}&\textbf{61.4}&\textbf{60.4}&\textbf{59.7}\\\\
CARP (mc)&200&78.7&79.7&35&26.6&\textbf{14.5}&61.8&95.5&64.7&57.1&57.1&55.9&55\\\\
&400&83.9&80.3&34.8&27.1&14.2&\textbf{62.9}&95.5&62.8&57.6&57.7&56.8&56\\\\
\hline
\end{array}