Noisy data, crawled from the web or supplied by volunteers such as Mechanical\nTurkers or citizen scientists, is considered an alternative to professionally\nlabeled data. There has been research focused on mitigating the effects of\nlabel noise. It is typically modeled as inaccuracy, where the correct label is\nreplaced by an incorrect label from the same set. We consider an additional\ndimension of label noise: imprecision. For example, a non-breeding snow bunting\nis labeled as a bird. This label is correct, but not as precise as the task\nrequires.\n Standard softmax classifiers cannot learn from such a weak label because they\nconsider all classes mutually exclusive, which non-breeding snow bunting and\nbird are not. We propose CHILLAX (Class Hierarchies for Imprecise Label\nLearning and Annotation eXtrapolation), a method based on hierarchical\nclassification, to fully utilize labels of any precision.\n Experiments on noisy variants of NABirds and ILSVRC2012 show that our method\noutperforms strong baselines by as much as 16.4 percentage points, and the\ncurrent state of the art by up to 3.9 percentage points.\n
Paper
References (33)
Scroll for more · 21 remaining