We present a benchmark suite of four datasets for evaluating the fairness of\npre-trained language models and the techniques used to fine-tune them for\ndownstream tasks. Our benchmarks cover four jurisdictions (European Council,\nUSA, Switzerland, and China), five languages (English, German, French, Italian\nand Chinese) and fairness across five attributes (gender, age, region,\nlanguage, and legal area). In our experiments, we evaluate pre-trained language\nmodels using several group-robust fine-tuning techniques and show that\nperformance group disparities are vibrant in many cases, while none of these\ntechniques guarantee fairness, nor consistently mitigate group disparities.\nFurthermore, we provide a quantitative and qualitative analysis of our results,\nhighlighting open challenges in the development of robustness methods in legal\nNLP.\n
Paper
References (66)
Scroll for more · 38 remaining