Multilingual Twitter Corpus and Baselines for Evaluating Demographic Bias in Hate Speech Recognition
Existing research on fairness evaluation of document classification models\nmainly uses synthetic monolingual data without ground truth for author\ndemographic attributes. In this work, we assemble and publish a multilingual\nTwitter corpus for the task of hate speech detection with inferred four author\ndemographic factors: age, country, gender and race/ethnicity. The corpus covers\nfive languages: English, Italian, Polish, Portuguese and Spanish. We evaluate\nthe inferred demographic labels with a crowdsourcing platform, Figure Eight. To\nexamine factors that can cause biases, we take an empirical analysis of\ndemographic predictability on the English corpus. We measure the performance of\nfour popular document classifiers and evaluate the fairness and bias of the\nbaseline classifiers on the author-level demographic attributes.\n