Pofcde: A Prompt-Based Debias Framework with Counterfactual Data Expansion

In recent years, pre-trained language models (PLMs) have shown remarkable performance in natural language processing tasks. However, they inevitably inherit biases from their training data. These biases can manifest as stereotypes or unfair predictions in real-world applications, potentially leading to serious issues, especially in sensitive contexts. While many existing debiasing methods can mitigate such problems to some extent, they come with notable limitations. Specifically, when models are fine-tuned for downstream tasks, newly introduced biases often resurface, undermining the long-term effectiveness of debiasing efforts and sometimes negatively impacting task performance. To address these challenges, we propose Pofcde, a debiasing framework designed to prevent biases from being reintroduced during task adaptation. Our approach combines prompt tuning and counterfactual contrastive learning to offer a flexible and efficient solution. Rather than modifying model parameters directly, we use trainable prompt vectors to guide the model's adaptation to downstream tasks, preserving the knowledge acquired during pre-training. Additionally, we employ a counterfactual data augmentation strategy to create pairs of sentences that share similar meanings but differ in bias direction. This enables the model to learn how to make stable and consistent predictions even when bias-related features vary. To further enhance data diversity, we incorporate external corpora, expanding the model's exposure to a wider range of scenarios. Experimental results demonstrate that Pofcde effectively prevents the reintroduction of bias during fine-tuning, ensuring both fairness and robust task performance.

Paper

Full text

PDF

Pofcde: A Prompt-Based Debias Framework with Counterfactual Data Expansion

Semantic Scholar · 2025

Abstract

In recent years, pre-trained language models (PLMs) have shown remarkable performance in natural language processing tasks. However, they inevitably inherit biases from their training data. These biases can manifest as stereotypes or unfair predictions in real-world applications, potentially leading to serious issues, especially in sensitive contexts. While many existing debiasing methods can mitigate such problems to some extent, they come with notable limitations. Specifically, when models are fine-tuned for downstream tasks, newly introduced biases often resurface, undermining the long-term effectiveness of debiasing efforts and sometimes negatively impacting task performance. To address these challenges, we propose Pofcde, a debiasing framework designed to prevent biases from being reintroduced during task adaptation. Our approach combines prompt tuning and counterfactual contrastive learning to offer a flexible and efficient solution. Rather than modifying model parameters directly, we use trainable prompt vectors to guide the model's adaptation to downstream tasks, preserving the knowledge acquired during pre-training. Additionally, we employ a counterfactual data augmentation strategy to create pairs of sentences that share similar meanings but differ in bias direction. This enables the model to learn how to make stable and consistent predictions even when bias-related features vary. To further enhance data diversity, we incorporate external corpora, expanding the model's exposure to a wider range of scenarios. Experimental results demonstrate that Pofcde effectively prevents the reintroduction of bias during fine-tuning, ensuring both fairness and robust task performance.

Similar papers

© 2026 NYSGPT2525 LLC