> It would be beneficial to provide a more detailed explanation of the rationale behind the selection of specific search types, optimizers, and supernet types for the comparative evaluation.
In the single-stage neural architecture search literature [1], it is well-known that several of the single-stage optimizers such as DARTS and GDAS have several failure modes. We started our study on weight entanglement spaces by studying DrNAS, DARTS and GDAS on the two tiny search spaces which we create benchmarks for. We find in our initial analysis that DrNAS performs competitively and is robust on both of these spaces in comparison to DARTS and GDAS. Hence we restrict ourselves to DrNAS in our large scale experiments. We have also evaluated DARTS and GDAS on the AutoFormer-T space on CIFAR-100 and presented the results below.
Toy conv macro (CIFAR10)
| Optimizer | Test-acc |
|------------|-------------------|
| DrNAS | **83.02** |
| SPOS+RS | 81.2525 |
| SPOS+RE | 81.89 |
| DARTS_v1 | 81.61 |
| DARTS_v2 | 81.49 |
| GDAS | 10% (degenerate) |
Toy cell entangled (Fashion MNIST)
| Optimizer | Test-acc |
|------------|-----------|
| DrNAS | **90.93** |
| DARTS-V1 | 89.905 |
| DARTS-V2 | 90.7475 |
| GDAS | 90.618 |
| SPOS+RS | 90.6875 |
| SPOS+RE | 90.595 |
AutoFormer-T (CIFAR100)
| Optimizer | CIFAR100 Test-acc |
|------------|---------------|
|TangleNAS-DARTS | 82.107 ± 0.392 |
| TangleNAS-GDAS | 82.12 ± 0.2813 |
| TangleNAS-DrNAS | **82.668 ± 0.161** |
| SPOS+ES | 82.5175 ± 0.114 |
| SPOS+RS | 82.210 ± 0.14242 |
We choose our search spaces with three primary factors in mind (1) Architecture Diversity: ViTs, LLMs, mobilenets i.e. convolutional space (2) Application Diversity: Classification and Language Modelling) (3) Usefulness: A focus on spaces constructed around foundation model architectures
[1] White, C., Safari, M., Sukthanker, R., Ru, B., Elsken, T., Zela, A., Dey, D. and Hutter, F., 2023. Neural architecture search: Insights from 1000 papers. arXiv preprint arXiv:2301.08727.
> To discuss potential limitations ...
We thank the reviewer for this question. In our opinion the primary limitation of our work is the inability to handle multiple architecture constraints like parameter-sizes, FLOPS directly. It is indeed possible to add a constrained secondary penalty to our loss. However , this modification still doesn’t allow one to generate a full Pareto Front of objectives which is useful in many deep learning applications (eg: fairness, robustness, hardware constraints). We leave this extension for future work. Further, like 2-stage methods (OFA, AutoFormer), our approach also requires some amount of engineering effort to shard and combine weights into a mixture.
Furthermore now we also conduct an experiment to search for models with smaller parameter sizes. We add a differentiable parameter penalty to the loss function to enforce this. We then perform the evolutionary search of SPOS setting the parameter size of the TangleNAS model as the constraint. We find that the architecture SPOS discovers in this constrained version is still worse than the TangleNAS architecture discovered. This shows the ability of our search method to discover better architectures even at smaller parameter budgets in its constrained version.
| Optimizer | CIFAR10 | Params | CIFAR100 | Params |
|-----------|-------------------|-----------|------------------|---------|
| TangleNAS | **97.254925±0.11523** | 6.647626M | **81.25467±0.27151** | 7.08394M |
| SPOS+RE | 97.16425±0.13604 | 6.56041M | 80.8324±0.23757 | 6.8866M |
If the reviewer has any follow-up questions we would be happy to address them.