Does the Objective Matter? Comparing Training Objectives for Pronoun Resolution

Hard cases of pronoun resolution have been used as a long-standing benchmark\nfor commonsense reasoning. In the recent literature, pre-trained language\nmodels have been used to obtain state-of-the-art results on pronoun resolution.\nOverall, four categories of training and evaluation objectives have been\nintroduced. The variety of training datasets and pre-trained language models\nused in these works makes it unclear whether the choice of training objective\nis critical. In this work, we make a fair comparison of the performance and\nseed-wise stability of four models that represent the four categories of\nobjectives. Our experiments show that the objective of sequence ranking\nperforms the best in-domain, while the objective of semantic similarity between\ncandidates and pronoun performs the best out-of-domain. We also observe a\nseed-wise instability of the model using sequence ranking, which is not the\ncase when the other objectives are used.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC