Models based on the transformer architecture, such as BERT, have marked a\ncrucial step forward in the field of Natural Language Processing. Importantly,\nthey allow the creation of word embeddings that capture important semantic\ninformation about words in context. However, as single entities, these\nembeddings are difficult to interpret and the models used to create them have\nbeen described as opaque. Binder and colleagues proposed an intuitive embedding\nspace where each dimension is based on one of 65 core semantic features.\nUnfortunately, the space only exists for a small dataset of 535 words, limiting\nits uses. Previous work (Utsumi, 2018, 2020, Turton, Vinson & Smith, 2020) has\nshown that Binder features can be derived from static embeddings and\nsuccessfully extrapolated to a large new vocabulary. Taking the next step, this\npaper demonstrates that Binder features can be derived from the BERT embedding\nspace. This provides contextualised Binder embeddings, which can aid in\nunderstanding semantic differences between words in context. It additionally\nprovides insights into how semantic features are represented across the\ndifferent layers of the BERT model.\n
Paper
References (37)
Scroll for more · 25 remaining