Information retrieval for label noise document ranking by bag sampling and group-wise loss

Long Document retrieval (DR) has always been a tremendous challenge for\nreading comprehension and information retrieval. The pre-training model has\nachieved good results in the retrieval stage and Ranking for long documents in\nrecent years. However, there is still some crucial problem in long document\nranking, such as data label noises, long document representations, negative\ndata Unbalanced sampling, etc. To eliminate the noise of labeled data and to be\nable to sample the long documents in the search reasonably negatively, we\npropose the bag sampling method and the group-wise Localized Contrastive\nEstimation(LCE) method. We use the head middle tail passage for the long\ndocument to encode the long document, and in the retrieval, stage Use dense\nretrieval to generate the candidate's data. The retrieval data is divided into\nmultiple bags at the ranking stage, and negative samples are selected in each\nbag. After sampling, two losses are combined. The first loss is LCE. To fit bag\nsampling well, after query and document are encoded, the global features of\neach group are extracted by convolutional layer and max-pooling to improve the\nmodel's resistance to the impact of labeling noise, finally, calculate the LCE\ngroup-wise loss. Notably, our model shows excellent performance on the MS MARCO\nLong document ranking leaderboard.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC