Type- and token-based embedding architectures are still competing in lexical\nsemantic change detection. The recent success of type-based models in\nSemEval-2020 Task 1 has raised the question why the success of token-based\nmodels on a variety of other NLP tasks does not translate to our field. We\ninvestigate the influence of a range of variables on clusterings of BERT\nvectors and show that its low performance is largely due to orthographic\ninformation on the target word, which is encoded even in the higher layers of\nBERT representations. By reducing the influence of orthography we considerably\nimprove BERT's performance.\n
Paper
References (33)
Scroll for more · 21 remaining