Exploring Opportunistic Meta-knowledge to Reduce Search Spaces for Automated Machine Learning

Machine learning (ML) pipeline composition and optimisation have been studied\nto seek multi-stage ML models, i.e. preprocessor-inclusive, that are both valid\nand well-performing. These processes typically require the design and traversal\nof complex configuration spaces consisting of not just individual ML components\nand their hyperparameters, but also higher-level pipeline structures that link\nthese components together. Optimisation efficiency and resulting ML-model\naccuracy both suffer if this pipeline search space is unwieldy and excessively\nlarge; it becomes an appealing notion to avoid costly evaluations of poorly\nperforming ML components ahead of time. Accordingly, this paper investigates\nwhether, based on previous experience, a pool of available\nclassifiers/regressors can be preemptively culled ahead of initiating a\npipeline composition/optimisation process for a new ML problem, i.e. dataset.\nThe previous experience comes in the form of classifier/regressor accuracy\nrankings derived, with loose assumptions, from a substantial but non-exhaustive\nnumber of pipeline evaluations; this meta-knowledge is considered\n'opportunistic'. Numerous experiments with the AutoWeka4MCPS package, including\nones leveraging similarities between datasets via the relative landmarking\nmethod, show that, despite its seeming unreliability, opportunistic\nmeta-knowledge can improve ML outcomes. However, results also indicate that the\nculling of classifiers/regressors should not be too severe either. In effect,\nit is better to search through a 'top tier' of recommended predictors than to\npin hopes onto one previously supreme performer.\n

Paper

References (21)

Scroll for more · 9 remaining

Similar papers

© 2026 NYSGPT2525 LLC