In this paper we introduce two algorithms for neural architecture search\n(NASGD and NASAGD) following the theoretical work by two of the authors [5]\nwhich used the geometric structure of optimal transport to introduce the\nconceptual basis for new notions of traditional and accelerated gradient\ndescent algorithms for the optimization of a function on a semi-discrete space.\nOur algorithms, which use the network morphism framework introduced in [2] as a\nbaseline, can analyze forty times as many architectures as the hill climbing\nmethods [2, 14] while using the same computational resources and time and\nachieving comparable levels of accuracy. For example, using NASGD on CIFAR-10,\nour method designs and trains networks with an error rate of 4.06 in only 12\nhours on a single GPU.\n