Teacher-Student Training and Triplet Loss to Reduce the Effect of Drastic Face Occlusion

We study a series of recognition tasks in two realistic scenarios requiring\nthe analysis of faces under strong occlusion. On the one hand, we aim to\nrecognize facial expressions of people wearing Virtual Reality (VR) headsets.\nOn the other hand, we aim to estimate the age and identify the gender of people\nwearing surgical masks. For all these tasks, the common ground is that half of\nthe face is occluded. In this challenging setting, we show that convolutional\nneural networks (CNNs) trained on fully-visible faces exhibit very low\nperformance levels. While fine-tuning the deep learning models on occluded\nfaces is extremely useful, we show that additional performance gains can be\nobtained by distilling knowledge from models trained on fully-visible faces. To\nthis end, we study two knowledge distillation methods, one based on\nteacher-student training and one based on triplet loss. Our main contribution\nconsists in a novel approach for knowledge distillation based on triplet loss,\nwhich generalizes across models and tasks. Furthermore, we consider combining\ndistilled models learned through conventional teacher-student training or\nthrough our novel teacher-student training based on triplet loss. We provide\nempirical evidence showing that, in most cases, both individual and combined\nknowledge distillation methods bring statistically significant performance\nimprovements. We conduct experiments with three different neural models (VGG-f,\nVGG-face, ResNet-50) on various tasks (facial expression recognition, gender\nrecognition, age estimation), showing consistent improvements regardless of the\nmodel or task.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC