Towards A Friendly Online Community: An Unsupervised Style Transfer Framework for Profanity Redaction
Offensive and abusive language is a pressing problem on social media\nplatforms. In this work, we propose a method for transforming offensive\ncomments, statements containing profanity or offensive language, into\nnon-offensive ones. We design a RETRIEVE, GENERATE and EDIT unsupervised style\ntransfer pipeline to redact the offensive comments in a word-restricted manner\nwhile maintaining a high level of fluency and preserving the content of the\noriginal text. We extensively evaluate our method's performance and compare it\nto previous style transfer models using both automatic metrics and human\nevaluations. Experimental results show that our method outperforms other models\non human evaluations and is the only approach that consistently performs well\non all automatic evaluation metrics.\n