In pickup and delivery services, transaction classification based on customer\nprovided free text is a challenging problem. It involves the association of a\nwide variety of customer inputs to a fixed set of categories while adapting to\nthe various customer writing styles. This categorization is important for the\nbusiness: it helps understand the market needs and trends, and also assist in\nbuilding a personalized experience for different segments of the customers.\nHence, it is vital to capture these category information trends at scale, with\nhigh precision and recall. In this paper, we focus on a specific use-case where\na single category drives each transaction. We propose a cost-effective\ntransaction classification approach based on semi-supervision and knowledge\ndistillation frameworks. The approach identifies the category of a transaction\nusing free text input given by the customer. We use weak labelling and notice\nthat the performance gains are similar to that of using human-annotated\nsamples. On a large internal dataset and on 20Newsgroup dataset, we see that\nRoBERTa performs the best for the categorization tasks. Further, using an\nALBERT model (it has 33x fewer parameters vis-a-vis parameters of RoBERTa),\nwith RoBERTa as the Teacher, we see a performance similar to that of RoBERTa\nand better performance over unadapted ALBERT. This framework, with ALBERT as a\nstudent and RoBERTa as teacher, is further referred to as R-ALBERT in this\npaper. The model is in production and is used by business to understand\nchanging trends and take appropriate decisions.\n
Paper
References (35)
Scroll for more · 23 remaining