Interpretability Analysis for Named Entity Recognition to Understand System Predictions and How They Can Improve
Named Entity Recognition systems achieve remarkable performance on domains\nsuch as English news. It is natural to ask: What are these models actually\nlearning to achieve this? Are they merely memorizing the names themselves? Or\nare they capable of interpreting the text and inferring the correct entity type\nfrom the linguistic context? We examine these questions by contrasting the\nperformance of several variants of LSTM-CRF architectures for named entity\nrecognition, with some provided only representations of the context as\nfeatures. We also perform similar experiments for BERT. We find that context\nrepresentations do contribute to system performance, but that the main factor\ndriving high performance is learning the name tokens themselves. We enlist\nhuman annotators to evaluate the feasibility of inferring entity types from the\ncontext alone and find that, while people are not able to infer the entity type\neither for the majority of the errors made by the context-only system, there is\nsome room for improvement. A system should be able to recognize any name in a\npredictive context correctly and our experiments indicate that current systems\nmay be further improved by such capability.\n
Paper
References (36)
Scroll for more · 24 remaining