Using a meta-model to compensate for training-evaluation mismatches

One of the fundamental assumptions of machine learning is
\nthat learnt models are applied to data that is identically distributed to
\nthe training data. This assumption is often not realistic: for example,
\ndata collected from a single source at different times may not be distributed identically, due to sampling bias or changes in the environment.
\nWe propose a new architecture called a meta-model which predicts performance for unseen models. This approach is applicable when several
\n‘proxy’ datasets are available to train a model to be deployed on a ‘target’
\ntest set; the architecture is used to identify which regression algorithms
\nshould be used as well as which datasets are most useful to train for a
\ngiven target dataset. Finally, we demonstrate the strengths and weaknesses of the proposed meta-model by making use of artificially generated
\ndatasets using a variation of the Friedman method 3 used to generate
\nartificial regression datasets, and discuss real-world applications of our
\napproach.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC