Rethinking AI Safety and Ethics: A State-of-the-Art Multi-Task Model for Bias and Toxicity Detection via Task-Specific Supervision and Data-centric fine-tuning
Advancements in large language models (LLMs) have massively transformed user interactions with AI by enabling the creation of systems that exhibit human-like capabilities, particularly in natural language processing and understanding, allowing for more natural and intuitive human-computer interactions. While the transformative potential of LLMs is undeniable, these powerful models also come with potential risks of generating biased, inappropriate and harmful responses, which needs to be mitigated; and in this process of risk management of LLMs, a very critical step is the detection of these risks.The existing models for toxicity and bias detection, primarily been trained on datasets sourced from social media or similar public platforms, have significant performance gaps when applied to real-world user-AI interactions [1]. To address these gaps, we have introduced a multi-task T5-based model [2] specifically designed for simultaneously detecting toxicity and bias within conversational AI. Our approach involved an iterative, three-stage methodology, which involves synthetic training dataset creation, balancing the dataset through strategic downsampling of dominant categories and upsampling of minority categories, and iterative model training mechanism to accommodate computational limitations. Our research emphasizes a data-centric fine-tuning approach, the importance of carefully curated, balanced training datasets, iterative model refinement, and comprehensive evaluation methodologies. Evaluation on benchmark datasets demonstrated our model’s superior performance, exceeding Claude V2 [3] and achieving comparable accuracy to GPT 4 [4] and superior accuracy to GPT-3.5 [5] in a typical setting of LLM-as-a-judge based evaluation in Phoenix evaluation framework [6]. Our work provides a foundational framework towards building safer, ethically aligned and explainable conversational AI systems.
Paper
Full text
Rethinking AI Safety and Ethics: A State-of-the-Art Multi-Task Model for Bias and Toxicity Detection via Task-Specific Supervision and Data-centric fine-tuning
Semantic Scholar · 2025
Abstract
Advancements in large language models (LLMs) have massively transformed user interactions with AI by enabling the creation of systems that exhibit human-like capabilities, particularly in natural language processing and understanding, allowing for more natural and intuitive human-computer interactions. While the transformative potential of LLMs is undeniable, these powerful models also come with potential risks of generating biased, inappropriate and harmful responses, which needs to be mitigated; and in this process of risk management of LLMs, a very critical step is the detection of these risks.The existing models for toxicity and bias detection, primarily been trained on datasets sourced from social media or similar public platforms, have significant performance gaps when applied to real-world user-AI interactions [1]. To address these gaps, we have introduced a multi-task T5-based model [2] specifically designed for simultaneously detecting toxicity and bias within conversational AI. Our approach involved an iterative, three-stage methodology, which involves synthetic training dataset creation, balancing the dataset through strategic downsampling of dominant categories and upsampling of minority categories, and iterative model training mechanism to accommodate computational limitations. Our research emphasizes a data-centric fine-tuning approach, the importance of carefully curated, balanced training datasets, iterative model refinement, and comprehensive evaluation methodologies. Evaluation on benchmark datasets demonstrated our model’s superior performance, exceeding Claude V2 [3] and achieving comparable accuracy to GPT 4 [4] and superior accuracy to GPT-3.5 [5] in a typical setting of LLM-as-a-judge based evaluation in Phoenix evaluation framework [6]. Our work provides a foundational framework towards building safer, ethically aligned and explainable conversational AI systems.