Landmark-Aware and Part-based Ensemble Transfer Learning Network for Facial Expression Recognition from Static images
Facial Expression Recognition from static images is a challenging problem in\ncomputer vision applications. Convolutional Neural Network (CNN), the\nstate-of-the-art method for various computer vision tasks, has had limited\nsuccess in predicting expressions from faces having extreme poses,\nillumination, and occlusion conditions. To mitigate this issue, CNNs are often\naccompanied by techniques like transfer, multi-task, or ensemble learning that\noften provide high accuracy at the cost of increased computational complexity.\nIn this work, we propose a Part-based Ensemble Transfer Learning network that\nmodels how humans recognize facial expressions by correlating the spatial\norientation pattern of the facial features with a specific expression. It\nconsists of 5 sub-networks, and each sub-network performs transfer learning\nfrom one of the five subsets of facial landmarks: eyebrows, eyes, nose, mouth,\nor jaw to expression classification. We show that our proposed ensemble network\nuses visual patterns emanating from facial muscles' motor movements to predict\nexpressions and demonstrate the usefulness of transfer learning from Facial\nLandmark Localization to Facial Expression Recognition. We test the proposed\nnetwork on the CK+, JAFFE, and SFEW datasets, and it outperforms the benchmark\nfor CK+ and JAFFE datasets by 0.51% and 5.34%, respectively. Additionally, the\nproposed ensemble network consists of only 1.65M model parameters, ensuring\ncomputational efficiency during training and real-time deployment. Grad-CAM\nvisualizations of our proposed ensemble highlight the complementary nature of\nits sub-networks, a key design parameter of an effective ensemble network.\nLastly, cross-dataset evaluation results reveal that our proposed ensemble has\na high generalization capacity, making it suitable for real-world usage.\n
Paper
References (62)
Scroll for more · 38 remaining