Bootstrapping the Out-of-sample Predictions for Efficient and Accurate Cross-Validation

Cross-Validation (CV), and out-of-sample performance-estimation protocols in\ngeneral, are often employed both for (a) selecting the optimal combination of\nalgorithms and values of hyper-parameters (called a configuration) for\nproducing the final predictive model, and (b) estimating the predictive\nperformance of the final model. However, the cross-validated performance of the\nbest configuration is optimistically biased. We present an efficient bootstrap\nmethod that corrects for the bias, called Bootstrap Bias Corrected CV (BBC-CV).\nBBC-CV's main idea is to bootstrap the whole process of selecting the\nbest-performing configuration on the out-of-sample predictions of each\nconfiguration, without additional training of models. In comparison to the\nalternatives, namely the nested cross-validation and a method by Tibshirani and\nTibshirani, BBC-CV is computationally more efficient, has smaller variance and\nbias, and is applicable to any metric of performance (accuracy, AUC,\nconcordance index, mean squared error). Subsequently, we employ again the idea\nof bootstrapping the out-of-sample predictions to speed up the CV process.\nSpecifically, using a bootstrap-based hypothesis test we stop training of\nmodels on new folds of statistically-significantly inferior configurations. We\nname the method Bootstrap Corrected with Early Dropping CV (BCED-CV) that is\nboth efficient and provides accurate performance estimates.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC