Data augmentation techniques are widely used for enhancing the performance of\nmachine learning models by tackling class imbalance issues and data sparsity.\nState-of-the-art generative language models have been shown to provide\nsignificant gains across different NLP tasks. However, their applicability to\ndata augmentation for text classification tasks in few-shot settings have not\nbeen fully explored, especially for specialised domains. In this paper, we\nleverage GPT-2 (Radford A et al, 2019) for generating artificial training\ninstances in order to improve classification performance. Our aim is to analyse\nthe impact the selection process of seed training examples have over the\nquality of GPT-generated samples and consequently the classifier performance.\nWe perform experiments with several seed selection strategies that, among\nothers, exploit class hierarchical structures and domain expert selection. Our\nresults show that fine-tuning GPT-2 in a handful of label instances leads to\nconsistent classification improvements and outperform competitive baselines.\nFinally, we show that guiding this process through domain expert selection can\nlead to further improvements, which opens up interesting research avenues for\ncombining generative models and active learning.\n
Paper
References (61)
Scroll for more · 38 remaining