Automating information extraction from legal documents and formalising them into a machine understandable format has long been an integral challenge to legal reasoning. Most approaches in the past consist of highly complex solutions that use annotated syntactic structures and grammar to distil rules. The current research trend is to utilise state-of-the-art natural language processing (NLP) approaches to automate these tasks, with minimum human interference. In this paper, based on its functional features, we propose a taxonomy of semantic type in korean legislation, such as obligations, rights, permissions, penalties, etc.. Based on this, we performed automatic classification of legal norms with a rule-based classifier using a manually labelled dataset formed by three korean acts, i.e., Insurance Business Act, Banking Act and Financial Holding Companies Act, of the Korean legislation ( \(n=1237\) ) and a performance of \(F_{1}=0.97\) was reached. In contrast, several supervised machine learning based classifiers were implemented and a performance of F-measure = 0.99 was achieved.
Paper
An open-access PDF is published at link.springer.com. 44B holds its address, not the file.
Open PDF