Recently research has started focusing on avoiding undesired effects that\ncome with content moderation, such as censorship and overblocking, when dealing\nwith hatred online. The core idea is to directly intervene in the discussion\nwith textual responses that are meant to counter the hate content and prevent\nit from further spreading. Accordingly, automation strategies, such as natural\nlanguage generation, are beginning to be investigated. Still, they suffer from\nthe lack of sufficient amount of quality data and tend to produce\ngeneric/repetitive responses. Being aware of the aforementioned limitations, we\npresent a study on how to collect responses to hate effectively, employing\nlarge scale unsupervised language models such as GPT-2 for the generation of\nsilver data, and the best annotation strategies/neural architectures that can\nbe used for data filtering before expert validation/post-editing.\n
Paper
References (53)
Scroll for more · 38 remaining