Automatic definition extraction from texts is an important task that has\nnumerous applications in several natural language processing fields such as\nsummarization, analysis of scientific texts, automatic taxonomy generation,\nontology generation, concept identification, and question answering. For\ndefinitions that are contained within a single sentence, this problem can be\nviewed as a binary classification of sentences into definitions and\nnon-definitions. In this paper, we focus on automatic detection of one-sentence\ndefinitions in mathematical texts, which are difficult to separate from\nsurrounding text. We experiment with several data representations, which\ninclude sentence syntactic structure and word embeddings, and apply deep\nlearning methods such as the Convolutional Neural Network (CNN) and the Long\nShort-Term Memory network (LSTM), in order to identify mathematical\ndefinitions. Our experiments demonstrate the superiority of CNN and its\ncombination with LSTM, when applied on the syntactically-enriched input\nrepresentation. We also present a new dataset for definition extraction from\nmathematical texts. We demonstrate that this dataset is beneficial for training\nsupervised models aimed at extraction of mathematical definitions. Our\nexperiments with different domains demonstrate that mathematical definitions\nrequire special treatment, and that using cross-domain learning is inefficient\nfor that task.\n
Paper
References (49)
Scroll for more · 37 remaining