Rethinking Zero-shot Video Classification: End-to-end Training for Realistic Applications

Trained on large datasets, deep learning (DL) can accurately classify videos\ninto hundreds of diverse classes. However, video data is expensive to annotate.\nZero-shot learning (ZSL) proposes one solution to this problem. ZSL trains a\nmodel once, and generalizes to new tasks whose classes are not present in the\ntraining dataset. We propose the first end-to-end algorithm for ZSL in video\nclassification. Our training procedure builds on insights from recent video\nclassification literature and uses a trainable 3D CNN to learn the visual\nfeatures. This is in contrast to previous video ZSL methods, which use\npretrained feature extractors. We also extend the current benchmarking\nparadigm: Previous techniques aim to make the test task unknown at training\ntime but fall short of this goal. We encourage domain shift across training and\ntest data and disallow tailoring a ZSL model to a specific test dataset. We\noutperform the state-of-the-art by a wide margin. Our code, evaluation\nprocedure and model weights are available at\ngithub.com/bbrattoli/ZeroShotVideoClassification.\n

Paper

References (67)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC