Deep neural networks (DNNs) with the flexibility to learn good top-layer representations have eclipsed shallow kernel methods without that flexibility. Here, we take inspiration from DNNs to develop the first non-Bayesian deep kernel method, the deep kernel machine. In addition, we develop a solver for the intermediate layer kernels in deep kernel machines that converges in around 10 steps, exploiting matrix solvers initially developed in the control theory literature. These are many times faster the usual gradient descent approach and generalise to arbitrary architectures. While deep kernel machines currently scale poorly in the number of datapoints, we believe that this can be rectified in future work, allowing deep kernel machines to form the basis of a new class of much more efficient deep nonlinear function approximators.
Paper
References (55)
Scroll for more · 38 remaining