Abstract Wepresentaprobabilisticviewpointtomultiplekernellearningunifyingwell-knownregu-larisedriskapproachesandrecentadvancesinapproximateBayesianinferencerelaxations.Theframeworkproposesageneralobjectivefunctionsuitableforregression,robustregres-sion and classification that is lower bound of the marginal likelihood and contains manyregularisedriskapproachesasspecialcases. Furthermore,wederiveanefficientandprov-ablyconvergentoptimisationalgorithm.Keywords: Multiple kernel learning, approximate Bayesian inference, double loop algo-rithms,Gaussianprocesses 1. Introduction Nonparametric kernel methods, cornerstones of machine learning today, can be seen fromdifferentangles: asregularisedriskminimisationinfunctionspaces(ScholkopfandSmola,2002), or as probabilistic Gaussian process methods (Rasmussen and Williams,2006). Inthesetechniques,thekernel(orequivalentlycovariance)functionencodesinterpolationchar-acteristics from observed to unseen points, and two basic statistical problems have to bemastered. First,alatentfunctionmustbepredictedwhichfitsdatawell,yetisassmoothaspossiblegiventhefixedkernel. Second,thekernelfunctionparametershavetobelearnedaswell,tobestsupportpredictionswhichareofprimaryinterest. Whilethefirstproblemissimplerandhasbeenaddressedmuchmorefrequentlysofar,thecentralroleoflearningthecovariancefunctioniswellacknowledged, andasubstantialnumberofmethodsfor“learn-ing the kernel”, “multiple kernel learning”, or “evidence maximisation” are available now.However, much of this work has firmly been associated with one of the “camps” (referredto as regularised risk and probabilistic in the sequel) with surprisingly little crosstalk oracknowledgmentsofpriorworkacrossthisboundary. Inthispaper,weclarifytherelation-ship between major regularised risk and probabilistic kernel learning techniques precisely,pointingoutadvantagesandpitfallsofeither,aswellasalgorithmicsimilaritiesleadingtonovelpowerfulalgorithms.Wedevelopacommonanalyticalandalgorithmicalframeworkencompassingapproachesfrombothcampsandprovideclearinsightsintotheoptimisationstructure. Eventhough,most of the optimisation is non convex, we show how to operate a provably convergent“almostNewton”methodnevertheless. Eachstepisnotmuchmoreexpensivethanagradient
Paper
References (26)
Scroll for more · 14 remaining