Extensive Error Analysis and a Learning-Based Evaluation of Medical Entity Recognition Systems to Approximate User Experience

When comparing entities extracted by a medical entity recognition system with\ngold standard annotations over a test set, two types of mismatches might occur,\nlabel mismatch or span mismatch. Here we focus on span mismatch and show that\nits severity can vary from a serious error to a fully acceptable entity\nextraction due to the subjectivity of span annotations. For a domain-specific\nBERT-based NER system, we showed that 25% of the errors have the same labels\nand overlapping span with gold standard entities. We collected expert judgement\nwhich shows more than 90% of these mismatches are accepted or partially\naccepted by the user. Using the training set of the NER system, we built a fast\nand lightweight entity classifier to approximate the user experience of such\nmismatches through accepting or rejecting them. The decisions made by this\nclassifier are used to calculate a learning-based F-score which is shown to be\na better approximation of a forgiving user's experience than the relaxed\nF-score. We demonstrated the results of applying the proposed evaluation metric\nfor a variety of deep learning medical entity recognition models trained with\ntwo datasets.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC