Improving BERT Pre-training with Hard Negative Pairs

In this paper, we ran various experiments on BERT's pre-training tasks and observed their impact on a language model's success on different downstream tasks such as masked word prediction, sentiment analysis, named entity recognition and text classification. Also, an improvement method called Hard Negative Pairs (HNP) is suggested to increase the success of the Same Sentence Prediction (SSP) task. The goal of HNP is to pick negative pairs that are more similar to the original sentence. After that, two additional improvements for SSP are proposed. One of these methods aims to create shorter sequences while the other one's main goal is selecting a splitting point other than the middle of the sentence. Experiments were performed for these three different suggestions and their different combinations. The results show that the SSP model that uses HNP and smaller sequence generation methods has an improvement over the original SSP. The longest trained model on 1 GB data gets closer results to the BERTurk, even though BERTurk trained with 35 times more data.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC