AutoTestForge: A Multidimensional Automated Testing Framework for Natural Language Processing Models
In recent years, the application of behavioral testing in Natural Language Processing (NLP) model evaluation has experienced substantial growth. However, existing methods are restricted by reliance on manual labor and the limited scope of capability assessment. To address these limitations, we introduce AutoTestForge, an automated and multidimensional testing framework for NLP models. Through the integration of Large Language Models (LLMs) to automatically generate test templates and instantiate them, manual involvement is significantly reduced. Additionally, a mechanism for validating test case labels based on differential testing is proposed, which makes use of a multi-model voting system to guarantee the quality of test cases. The framework expands the test suite across three dimensions: taxonomy, fairness, and robustness, offering a comprehensive evaluation of the capabilities of NLP models. This expansion enables a in-depth and thorough assessment of the models, providing valuable insights into their strengths and weaknesses. A comprehensive evaluation across the sentiment analysis (SA) task and semantic textual similarity (STS) task demonstrates that AutoTestForge consistently outperforms existing datasets and testing tools with higher failure rates (an average of \(32.35\%\) for SA and \(31.61\%\) for STS). Moreover, different generation strategies exhibit stable effectiveness with failure rates ranging from \(25.77\%-38.04\%\) .
Paper
References (51)
Scroll for more · 38 remaining