Generative Model-driven Structure Aligning Discriminative Embeddings for Transductive Zero-shot Learning

Zero-shot Learning (ZSL) is a transfer learning technique which aims at\ntransferring knowledge from seen classes to unseen classes. This knowledge\ntransfer is possible because of underlying semantic space which is common to\nseen and unseen classes. Most existing approaches learn a projection function\nusing labelled seen class data which maps visual data to semantic data. In this\nwork, we propose a shallow but effective neural network-based model for\nlearning such a projection function which aligns the visual and semantic data\nin the latent space while simultaneously making the latent space embeddings\ndiscriminative. As the above projection function is learned using the seen\nclass data, the so-called projection domain shift exists. We propose a\ntransductive approach to reduce the effect of domain shift, where we utilize\nunlabeled visual data from unseen classes to generate corresponding semantic\nfeatures for unseen class visual samples. While these semantic features are\ninitially generated using a conditional variational auto-encoder, they are used\nalong with the seen class data to improve the projection function. We\nexperiment on both inductive and transductive setting of ZSL and generalized\nZSL and show superior performance on standard benchmark datasets AWA1, AWA2,\nCUB, SUN, FLO, and APY. We also show the efficacy of our model in the case of\nextremely less labelled data regime on different datasets in the context of\nZSL.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC