Named Entity Recognition for Biomedical Patent Text using Bi-LSTM Variants

Recent years have shown a substantial increase in biomedical publications (patents or scientific articles) that are multiplying at a daily pace. This has led to an increased interest in the extraction of meaningful information (e.g., named entities) from these publications. Traditional NER approaches demand a considerable level of engineering skills and domain expertise in designing rules and features for better algorithm accuracy. In addition, due to the structure and linguistic complexity of the patent text, constructing such rules and features is often a challenging task. In this paper, we investigate various variants of the Bi-LSTM model performance for NER task based on features generated automatically from an unlabelled genes and proteins patent corpora. The proposed model is able to capture the context representation of an input sequence and globally assign the related labels for each token. The CHARS-Bi-LSTM-EMA variant yielded the best performance and significantly outperformed the state-of-the art approach.

Paper

Full text

PDF

Named Entity Recognition for Biomedical Patent Text using Bi-LSTM Variants

Semantic Scholar · Computer Science · 2019

Abstract

Recent years have shown a substantial increase in biomedical publications (patents or scientific articles) that are multiplying at a daily pace. This has led to an increased interest in the extraction of meaningful information (e.g., named entities) from these publications. Traditional NER approaches demand a considerable level of engineering skills and domain expertise in designing rules and features for better algorithm accuracy. In addition, due to the structure and linguistic complexity of the patent text, constructing such rules and features is often a challenging task. In this paper, we investigate various variants of the Bi-LSTM model performance for NER task based on features generated automatically from an unlabelled genes and proteins patent corpora. The proposed model is able to capture the context representation of an input sequence and globally assign the related labels for each token. The CHARS-Bi-LSTM-EMA variant yielded the best performance and significantly outperformed the state-of-the art approach.

Similar papers

© 2026 NYSGPT2525 LLC