RuREBus: a Case Study of Joint Named Entity Recognition and Relation Extraction from e-Government Domain

We show-case an application of information extraction methods, such as named\nentity recognition (NER) and relation extraction (RE) to a novel corpus,\nconsisting of documents, issued by a state agency. The main challenges of this\ncorpus are: 1) the annotation scheme differs greatly from the one used for the\ngeneral domain corpora, and 2) the documents are written in a language other\nthan English. Unlike expectations, the state-of-the-art transformer-based\nmodels show modest performance for both tasks, either when approached\nsequentially, or in an end-to-end fashion. Our experiments have demonstrated\nthat fine-tuning on a large unlabeled corpora does not automatically yield\nsignificant improvement and thus we may conclude that more sophisticated\nstrategies of leveraging unlabelled texts are demanded. In this paper, we\ndescribe the whole developed pipeline, starting from text annotation, baseline\ndevelopment, and designing a shared task in hopes of improving the baseline.\nEventually, we realize that the current NER and RE technologies are far from\nbeing mature and do not overcome so far challenges like ours.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC