Temporal Stochastic Softmax for 3D CNNs: An Application in Facial Expression Recognition

Training deep learning models for accurate spatiotemporal recognition of\nfacial expressions in videos requires significant computational resources. For\npractical reasons, 3D Convolutional Neural Networks (3D CNNs) are usually\ntrained with relatively short clips randomly extracted from videos. However,\nsuch uniform sampling is generally sub-optimal because equal importance is\nassigned to each temporal clip. In this paper, we present a strategy for\nefficient video-based training of 3D CNNs. It relies on softmax temporal\npooling and a weighted sampling mechanism to select the most relevant training\nclips. The proposed softmax strategy provides several advantages: a reduced\ncomputational complexity due to efficient clip sampling, and an improved\naccuracy since temporal weighting focuses on more relevant clips during both\ntraining and inference. Experimental results obtained with the proposed method\non several facial expression recognition benchmarks show the benefits of\nfocusing on more informative clips in training videos. In particular, our\napproach improves performance and computational cost by reducing the impact of\ninaccurate trimming and coarse annotation of videos, and heterogeneous\ndistribution of visual information across time.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC