Memorization vs. Generalization: Quantifying Data Leakage in NLP Performance Evaluation

Public datasets are often used to evaluate the efficacy and generalizability\nof state-of-the-art methods for many tasks in natural language processing\n(NLP). However, the presence of overlap between the train and test datasets can\nlead to inflated results, inadvertently evaluating the model's ability to\nmemorize and interpreting it as the ability to generalize. In addition, such\ndata sets may not provide an effective indicator of the performance of these\nmethods in real world scenarios. We identify leakage of training data into test\ndata on several publicly available datasets used to evaluate NLP tasks,\nincluding named entity recognition and relation extraction, and study them to\nassess the impact of that leakage on the model's ability to memorize versus\ngeneralize.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC