Few-shot learning aims to generalize unseen classes that appear during\ntesting but are unavailable during training. Prototypical networks incorporate\nfew-shot metric learning, by constructing a class prototype in the form of a\nmean vector of the embedded support points within a class. The performance of\nprototypical networks in extreme few-shot scenarios (like one-shot) degrades\ndrastically, mainly due to the desuetude of variations within the clusters\nwhile constructing prototypes. In this paper, we propose to replace the typical\nprototypical loss function with an Episodic Triplet Mining (ETM) technique. The\nconventional triplet selection leads to overfitting, because of all possible\ncombinations being used during training. We incorporate episodic training for\nmining the semi hard positive and the semi hard negative triplets to overcome\nthe overfitting. We also propose an adaptation to make use of unlabeled\ntraining samples for better modeling. Experimenting on two different audio\nprocessing tasks, namely speaker recognition and audio event detection; show\nimproved performances and hence the efficacy of ETM over the prototypical loss\nfunction and other meta-learning frameworks. Further, we show improved\nperformances when unlabeled training samples are used.\n
Paper
References (28)
Scroll for more · 16 remaining