Understanding Fairness of Gender Classification Algorithms Across Gender-Race Groups

Automated gender classification has important applications in many domains,\nsuch as demographic research, law enforcement, online advertising, as well as\nhuman-computer interaction. Recent research has questioned the fairness of this\ntechnology across gender and race. Specifically, the majority of the studies\nraised the concern of higher error rates of the face-based gender\nclassification system for darker-skinned people like African-American and for\nwomen. However, to date, the majority of existing studies were limited to\nAfrican-American and Caucasian only. The aim of this paper is to investigate\nthe differential performance of the gender classification algorithms across\ngender-race groups. To this aim, we investigate the impact of (a) architectural\ndifferences in the deep learning algorithms and (b) training set imbalance, as\na potential source of bias causing differential performance across gender and\nrace. Experimental investigations are conducted on two latest large-scale\npublicly available facial attribute datasets, namely, UTKFace and FairFace. The\nexperimental results suggested that the algorithms with architectural\ndifferences varied in performance with consistency towards specific gender-race\ngroups. For instance, for all the algorithms used, Black females (Black race in\ngeneral) always obtained the least accuracy rates. Middle Eastern males and\nLatino females obtained higher accuracy rates most of the time. Training set\nimbalance further widens the gap in the unequal accuracy rates across all\ngender-race groups. Further investigations using facial landmarks suggested\nthat facial morphological differences due to the bone structure influenced by\ngenetic and environmental factors could be the cause of the least performance\nof Black females and Black race, in general.\n

Paper

References (26)

Scroll for more · 14 remaining

Similar papers

© 2026 NYSGPT2525 LLC