Confident Learning detects label errors using a classifier trained on the same noisy labels it audits, coupling the auditor to the data it checks. Estimating P(y|x) from feature-space geometry breaks this coupling, and SimiFeat established such training-free detection with local k-nearest-neighbour posteriors. We ask a narrower question: does pooling the geometry into bagged cluster posteriors (a "forest-of-means") beat local k-NN? A prerequisite observation shapes the comparison. The detector that turns a posterior into a flag matters independently of the estimator: for geometric posteriors a per-class voting rule beats Confident-Learning pruning, so any estimator comparison must hold the detector fixed. SimiFeat is stronger overall, but pooling overtakes it at high (>=40%) and feature-dependent noise, where a local neighbourhood is itself corrupted while a pooled region of hundreds of points is not; we confirm this crossover on real human noise (CIFAR-10N, ~40%: F1 0.930 vs. 0.915). The pooled estimator is thus a complement to SimiFeat scoped to high, feature-dependent noise, not a replacement.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex