Neural Tangent Kernel (NTK) theory is widely used to study the dynamics of\ninfinitely-wide deep neural networks (DNNs) under gradient descent. But do the\nresults for infinitely-wide networks give us hints about the behavior of real\nfinite-width ones? In this paper, we study empirically when NTK theory is valid\nin practice for fully-connected ReLU and sigmoid DNNs. We find out that whether\na network is in the NTK regime depends on the hyperparameters of random\ninitialization and the network's depth. In particular, NTK theory does not\nexplain the behavior of sufficiently deep networks initialized so that their\ngradients explode as they propagate through the network's layers: the kernel is\nrandom at initialization and changes significantly during training in this\ncase, contrary to NTK theory. On the other hand, in the case of vanishing\ngradients, DNNs are in the the NTK regime but become untrainable rapidly with\ndepth. We also describe a framework to study generalization properties of DNNs,\nin particular the variance of network's output function, by means of NTK theory\nand discuss its limits.\n