Detecting Offensive Content in Open-domain Conversations using Two Stage Semi-supervision

As open-ended human-chatbot interaction becomes commonplace, sensitive\ncontent detection gains importance. In this work, we propose a two stage\nsemi-supervised approach to bootstrap large-scale data for automatic sensitive\nlanguage detection from publicly available web resources. We explore various\ndata selection methods including 1) using a blacklist to rank online discussion\nforums by the level of their sensitiveness followed by randomly sampling\nutterances and 2) training a weakly supervised model in conjunction with the\nblacklist for scoring sentences from online discussion forums to curate a\ndataset. Our data collection strategy is flexible and allows the models to\ndetect implicit sensitive content for which manual annotations may be\ndifficult. We train models using publicly available annotated datasets as well\nas using the proposed large-scale semi-supervised datasets. We evaluate the\nperformance of all the models on Twitter and Toxic Wikipedia comments testsets\nas well as on a manually annotated spoken language dataset collected during a\nlarge scale chatbot competition. Results show that a model trained on this\ncollected data outperforms the baseline models by a large margin on both\nin-domain and out-of-domain testsets, achieving an F1 score of 95.5% on an\nout-of-domain testset compared to a score of 75% for models trained on public\ndatasets. We also showcase that large scale two stage semi-supervision\ngeneralizes well across multiple classes of sensitivities such as hate speech,\nracism, sexual and pornographic content, etc. without even providing explicit\nlabels for these classes, leading to an average recall of 95.5% versus the\nmodels trained using annotated public datasets which achieve an average recall\nof 73.2% across seven sensitive classes on out-of-domain testsets.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC