One of the fundamental assumptions of machine learning is \nthat learnt models are applied to data that is identically distributed to \nthe training data. This assumption is often not realistic: for example, \ndata collected from a single source at different times may not be distributed identically, due to sampling bias or changes in the environment. \nWe propose a new architecture called a meta-model which predicts performance for unseen models. This approach is applicable when several \n‘proxy’ datasets are available to train a model to be deployed on a ‘target’ \ntest set; the architecture is used to identify which regression algorithms \nshould be used as well as which datasets are most useful to train for a \ngiven target dataset. Finally, we demonstrate the strengths and weaknesses of the proposed meta-model by making use of artificially generated \ndatasets using a variation of the Friedman method 3 used to generate \nartificial regression datasets, and discuss real-world applications of our \napproach.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex