Large width limits have been a recent focus of deep learning research: modulo\ncomputational practicalities, do wider networks outperform narrower ones?\nAnswering this question has been challenging, as conventional networks gain\nrepresentational power with width, potentially masking any negative effects.\nOur analysis in this paper decouples capacity and width via the generalization\nof neural networks to Deep Gaussian Processes (Deep GP), a class of\nnonparametric hierarchical models that subsume neural nets. In doing so, we aim\nto understand how width affects (standard) neural networks once they have\nsufficient capacity for a given modeling task. Our theoretical and empirical\nresults on Deep GP suggest that large width can be detrimental to hierarchical\nmodels. Surprisingly, we prove that even nonparametric Deep GP converge to\nGaussian processes, effectively becoming shallower without any increase in\nrepresentational power. The posterior, which corresponds to a mixture of\ndata-adaptable basis functions, becomes less data-dependent with width. Our\ntail analysis demonstrates that width and depth have opposite effects: depth\naccentuates a model's non-Gaussianity, while width makes models increasingly\nGaussian. We find there is a "sweet spot" that maximizes test performance\nbefore the limiting GP behavior prevents adaptability, occurring at width = 1\nor width = 2 for nonparametric Deep GP. These results make strong predictions\nabout the same phenomenon in conventional neural networks trained with L2\nregularization (analogous to a Gaussian prior on parameters): we show that such\nneural networks may need up to 500 - 1000 hidden units for sufficient capacity\n- depending on the dataset - but further width degrades performance.\n