An Automated Framework for the Extraction of Semantic Legal Metadata\n from Legal Texts

Semantic legal metadata provides information that helps with understanding\nand interpreting legal provisions. Such metadata is therefore important for the\nsystematic analysis of legal requirements. However, manually enhancing a large\nlegal corpus with semantic metadata is prohibitively expensive. Our work is\nmotivated by two observations: (1) the existing requirements engineering (RE)\nliterature does not provide a harmonized view on the semantic metadata types\nthat are useful for legal requirements analysis; (2) automated support for the\nextraction of semantic legal metadata is scarce, and it does not exploit the\nfull potential of artificial intelligence technologies, notably natural\nlanguage processing (NLP) and machine learning (ML). Our objective is to take\nsteps toward overcoming these limitations. To do so, we review and reconcile\nthe semantic legal metadata types proposed in the RE literature. Subsequently,\nwe devise an automated extraction approach for the identified metadata types\nusing NLP and ML. We evaluate our approach through two case studies over the\nLuxembourgish legislation. Our results indicate a high accuracy in the\ngeneration of metadata annotations. In particular, in the two case studies, we\nwere able to obtain precision scores of 97.2% and 82.4% and recall scores of\n94.9% and 92.4%.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC