sigmoidF1: A Smooth F1 Score Surrogate Loss for Multilabel Classification

Multiclass multilabel classification is the task of attributing multiple\nlabels to examples via predictions. Current models formulate a reduction of the\nmultilabel setting into either multiple binary classifications or multiclass\nclassification, allowing for the use of existing loss functions (sigmoid,\ncross-entropy, logistic, etc.). Multilabel classification reductions do not\naccommodate for the prediction of varying numbers of labels per example and the\nunderlying losses are distant estimates of the performance metrics. We propose\na loss function, sigmoidF1, which is an approximation of the F1 score that (1)\nis smooth and tractable for stochastic gradient descent, (2) naturally\napproximates a multilabel metric, and (3) estimates label propensities and\nlabel counts. We show that any confusion matrix metric can be formulated with a\nsmooth surrogate. We evaluate the proposed loss function on text and image\ndatasets, and with a variety of metrics, to account for the complexity of\nmultilabel classification evaluation. sigmoidF1 outperforms other loss\nfunctions on one text and two image datasets and several metrics. These results\nshow the effectiveness of using inference-time metrics as loss functions for\nnon-trivial classification problems like multilabel classification.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC