Legal Document Retrieval using Document Vector Embeddings and Deep Learning

Domain specific information retrieval process has been a prominent and\nongoing research in the field of natural language processing. Many researchers\nhave incorporated different techniques to overcome the technical and domain\nspecificity and provide a mature model for various domains of interest. The\nmain bottleneck in these studies is the heavy coupling of domain experts, that\nmakes the entire process to be time consuming and cumbersome. In this study, we\nhave developed three novel models which are compared against a golden standard\ngenerated via the on line repositories provided, specifically for the legal\ndomain. The three different models incorporated vector space representations of\nthe legal domain, where document vector generation was done in two different\nmechanisms and as an ensemble of the above two. This study contains the\nresearch being carried out in the process of representing legal case documents\ninto different vector spaces, whilst incorporating semantic word measures and\nnatural language processing techniques. The ensemble model built in this study,\nshows a significantly higher accuracy level, which indeed proves the need for\nincorporation of domain specific semantic similarity measures into the\ninformation retrieval process. This study also shows, the impact of varying\ndistribution of the word similarity measures, against varying document vector\ndimensions, which can lead to improvements in the process of legal information\nretrieval.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC