Recent work raises concerns about the use of standard splits to compare\nnatural language processing models. We propose a Bayesian statistical model\ncomparison technique which uses k-fold cross-validation across multiple data\nsets to estimate the likelihood that one model will outperform the other, or\nthat the two will produce practically equivalent results. We use this technique\nto rank six English part-of-speech taggers across two data sets and three\nevaluation metrics.\n