Your fairness may vary: Pretrained language model fairness in toxic text classification

We study the performance-fairness trade-off in more than a dozen fine-tuned LMs for toxic text classification. We empirically show that no blanket statement can be made with respect to the bias of large versus regular versus compressed models. Moreover, we find that focusing on fairness-agnostic performance metrics can lead to models with varied fairness characteristics.

Paper

Similar papers

© 2026 NYSGPT2525 LLC