On the Trade-off of Intra-/Inter-class Diversity for Supervised Pre-training

Pre-training datasets are critical for building state-of-the-art machine learning models, motivating rigorous study on their impact on downstream tasks. In this work, we study the impact of the trade-off between the intra-class diversity (the number of samples per class) and the inter-class diversity (the number of classes) of a supervised pre-training dataset. Empirically, we found that with the size of the pre-training dataset fixed, the best downstream performance comes with a balance on the intra-/inter-class diversity. To understand the underlying mechanism, we show theoretically that the downstream performance depends monotonically on both types of diversity. Notably, our theory reveals that the optimal class-to-sample ratio (#classes / #samples per class) is invariant to the size of the pre-training dataset, which motivates an application of predicting the optimal number of pre-training classes. We demonstrate the effectiveness of this application by an improvement of around 2 points on the downstream tasks when using ImageNet as the pre-training dataset.

Paper

Similar papers

Peer review

Reviewer AXB64/10 · confidence 5/52023-07-05

Summary

This paper provides empirical study and theoretical analysis of the class-to-sample ratio in supervised pre-training datasets.

Strengths

1. This papers reveals an important conclusion - the optimal class-to-sample ratio is invariant to the size of the pre-training dataset. This can be utilized as a guidance of the data size for the scaling law in pre-training. 2. Detailed theoretical analysis are provided, from which a guidance of choosing the class-to-sample ratio is also provided.

Weaknesses

1. The pre-training dataset is restricted to ImageNet. 2. The model ResNet18 used is relatively small and may not reveal the true conclusion. 3. The evaluation is restricted to only 7 small downstream datasets.

Questions

1. Would other datasets affect the conclusion drawn. 2. Would other larger models and vision transformers lead to the similar observation and conclusion? 3. Would similar observation holds on larger downstream datasets or downstream datasets of other domains?

Rating

4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

See weakness and questions.

Reviewer Ypk96/10 · confidence 4/52023-07-06

Summary

The paper presents a detailed analysis on the intra-class diversity and inter-class diversity when you have a fixed dataset size. The authors develop some theoretical rule based on empirical results.

Strengths

- Provides a rigorous study on pretraining using ImageNet, that makes the empirical motivation clear for theoretical result. - Theory is relatively well-explained in a step-by-step manner. Proposes a method to balance intra-class and inter-class diversity, which leads to better performance on downstream tasks. - The paper is well-structured and easy to follow. - As we move towards a more data-centric view of deep learning, analysis of the composition of a dataset and how to best create a new dataset is a welcome area of research. This paper is one in this new area and sets a relatively decent standard for what should be done.

Weaknesses

- All of the experiments are based on using ImageNet as the pretraining dataset. This makes me uncertain of whether the result relates to ImageNet only, as there are particularities in relation to dataset acquisition, labelling, diversity between samples in a class etc. that makes ImageNet unique. Another pre-training dataset would be very welcome and I think perhaps a needed addition to this paper. - The paper seems like it is inspired by What Makes ImageNet Good for Transfer Learning by Huh et al, 2016; and this paper beyond just studying the trade-off between intra/inter-class diversity, actually ended generating a theoretically and empirically grounded rule for dataset construction.

Questions

- Why not add another pre-training dataset to the mix? - What happens when pre-training distribution doesn't match well with downstream dataset?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

- Analysis restricted to ImageNet pretraining.

Reviewer xyoX6/10 · confidence 2/52023-07-06

Summary

This work studied the impact of the trade-off between the intra-class diversity (the number of samples per class) and the inter-class diversity (the number of classes) of a supervised pre-training dataset. The authors found that with the size of the pre-training dataset fixed, the best downstream performance comes with a balance on the intra-/inter-class diversity and show that the downstream performance depends monotonically on both types of diversity theoretically.

Strengths

* The claims over how intra-class diversity and the inter-class diversity affect downstream tasks' performance were well supported by the empirical results, i.e., Fig. 1. * The authors provided a theory on the impact and verified the theoretical findings via empirical results.

Weaknesses

* The study was limited to fine-tuning on classification tasks only, other types of tasks, such as detection and segmentation, were not covered. * Though ResNet-18 is a classic architecture, it became less popular recently. It remains unclear if the same conclusion holds for other models (e.g., ViT) empirically.

Questions

* Can authors elaborate line 195: "obtain the label $y_i$ for each $x_i$ by performing clustering.", for example, which clustering method was used? what is the input/feature for clustering algorithm?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

N/A

Reviewer Ypk92023-08-18

Response

Q1/W1. Thank you. Much appreciated. Concern is resolved. W2. I think I listed this as a weakness in so far as it is work that was done before, however not in the depth here. I do agree that given the depth that the paper goes to it makes additional novel contributions. Q2. I think something more out of distribution would have been more appreciated (maybe like earth observation imagery or biological images), but I see that the results are decent in this slightly different domain. If the paper is published, I think this would make the paper more convincing. Based on rebuttal I will raise my score to a 6.

Authorsrebuttal2023-08-21

Thank you for raising the score!

We would like to thank the reviewer for the comment and for raising the score. We will make sure that the additional experiments are included in the revision.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC