We use a dataset of U.S. first names with labels based on predominant gender\nand racial group to examine the effect of training corpus frequency on\ntokenization, contextualization, similarity to initial representation, and bias\nin BERT, GPT-2, T5, and XLNet. We show that predominantly female and non-white\nnames are less frequent in the training corpora of these four language models.\nWe find that infrequent names are more self-similar across contexts, with\nSpearman's r between frequency and self-similarity as low as -.763. Infrequent\nnames are also less similar to initial representation, with Spearman's r\nbetween frequency and linear centered kernel alignment (CKA) similarity to\ninitial representation as high as .702. Moreover, we find Spearman's r between\nracial bias and name frequency in BERT of .492, indicating that lower-frequency\nminority group names are more associated with unpleasantness. Representations\nof infrequent names undergo more processing, but are more self-similar,\nindicating that models rely on less context-informed representations of\nuncommon and minority names which are overfit to a lower number of observed\ncontexts.\n
Paper
References (49)
Scroll for more · 37 remaining