Cross-domain Retrieval in the Legal and Patent Domains: a Reproducibility Study

Domain specific search has always been a challenging information retrieval\ntask due to several challenges such as the domain specific language, the unique\ntask setting, as well as the lack of accessible queries and corresponding\nrelevance judgements. In the last years, pretrained language models, such as\nBERT, revolutionized web and news search. Naturally, the community aims to\nadapt these advancements to cross-domain transfer of retrieval models for\ndomain specific search. In the context of legal document retrieval, Shao et al.\npropose the BERT-PLI framework by modeling the Paragraph Level Interactions\nwith the language model BERT. In this paper we reproduce the original\nexperiments, we clarify pre-processing steps, add missing scripts for framework\nsteps and investigate different evaluation approaches, however we are not able\nto reproduce the evaluation results. Contrary to the original paper, we\ndemonstrate that the domain specific paragraph-level modelling does not appear\nto help the performance of the BERT-PLI model compared to paragraph-level\nmodelling with the original BERT. In addition to our legal search\nreproducibility study, we investigate BERT-PLI for document retrieval in the\npatent domain. We find that the BERT-PLI model does not yet achieve performance\nimprovements for patent document retrieval compared to the BM25 baseline.\nFurthermore, we evaluate the BERT-PLI model for cross-domain retrieval between\nthe legal and patent domain on individual components, both on a paragraph and\ndocument-level. We find that the transfer of the BERT-PLI model on the\nparagraph-level leads to comparable results between both domains as well as\nfirst promising results for the cross-domain transfer on the document-level.\nFor reproducibility and transparency as well as to benefit the community we\nmake our source code and the trained models publicly available.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC