Abuse is Contextual, What about NLP? The Role of Context in Abusive Language Annotation and Detection
The datasets most widely used for abusive language detection contain lists of\nmessages, usually tweets, that have been manually judged as abusive or not by\none or more annotators, with the annotation performed at message level. In this\npaper, we investigate what happens when the hateful content of a message is\njudged also based on the context, given that messages are often ambiguous and\nneed to be interpreted in the context of occurrence. We first re-annotate part\nof a widely used dataset for abusive language detection in English in two\nconditions, i.e. with and without context. Then, we compare the performance of\nthree classification algorithms obtained on these two types of dataset, arguing\nthat a context-aware classification is more challenging but also more similar\nto a real application scenario.\n