Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets

Language models can generate harmful and biased outputs and exhibit\nundesirable behavior according to a given cultural context. We propose a\nProcess for Adapting Language Models to Society (PALMS) with Values-Targeted\nDatasets, an iterative process to significantly change model behavior by\ncrafting and fine-tuning on a dataset that reflects a predetermined set of\ntarget values. We evaluate our process using three metrics: quantitative\nmetrics with human evaluations that score output adherence to a target value,\ntoxicity scoring on outputs; and qualitative metrics analyzing the most common\nword associated with a given social category. Through each iteration, we add\nadditional training dataset examples based on observed shortcomings from\nevaluations. PALMS performs significantly better on all metrics compared to\nbaseline and control models for a broad range of GPT-3 language model sizes\nwithout compromising capability integrity. We find that the effectiveness of\nPALMS increases with model size. We show that significantly adjusting language\nmodel behavior is feasible with a small, hand-curated dataset.\n

Paper

References (41)

Scroll for more · 29 remaining

Similar papers

© 2026 NYSGPT2525 LLC