The 2nd stage accounts for the differences in both the geographic space (locations) and feature space (covariates and target variable). It uses cluster ensembles (CE) to split folds. First, all blocks acquired from the 1st stage are separately clustered based on locations, target variable, and covariates respectively. Then, as shown in figure 1, the CE is used to combine them together to reflect the differences in both the geographic and feature spaces. Geospatial prediction studies, such as soil mapping, ecological modeling, have extensively employed Machine Learning (ML) models. The evaluation of the model is a crucial step. To obtain reliable results, a test set that unbiasedly represents the prediction locations is needed. However, obtaining an additional test set is frequently impractical. Thus, the available samples must be partitioned into training and validation subsets to implement the evaluation.
Paper
Full text
Spatial+: A new cross-validation method to evaluate geospatial machine learning models
Semantic Scholar · Computer Science · 2023
Abstract
The 2nd stage accounts for the differences in both the geographic space (locations) and feature space (covariates and target variable). It uses cluster ensembles (CE) to split folds. First, all blocks acquired from the 1st stage are separately clustered based on locations, target variable, and covariates respectively. Then, as shown in figure 1, the CE is used to combine them together to reflect the differences in both the geographic and feature spaces. Geospatial prediction studies, such as soil mapping, ecological modeling, have extensively employed Machine Learning (ML) models. The evaluation of the model is a crucial step. To obtain reliable results, a test set that unbiasedly represents the prediction locations is needed. However, obtaining an additional test set is frequently impractical. Thus, the available samples must be partitioned into training and validation subsets to implement the evaluation.
References (52)
Scroll for more · 38 remaining