Benchmarking the Robustness of Spatial-Temporal Models Against Corruptions

The state-of-the-art deep neural networks are vulnerable to common\ncorruptions (e.g., input data degradations, distortions, and disturbances\ncaused by weather changes, system error, and processing). While much progress\nhas been made in analyzing and improving the robustness of models in image\nunderstanding, the robustness in video understanding is largely unexplored. In\nthis paper, we establish a corruption robustness benchmark, Mini Kinetics-C and\nMini SSV2-C, which considers temporal corruptions beyond spatial corruptions in\nimages. We make the first attempt to conduct an exhaustive study on the\ncorruption robustness of established CNN-based and Transformer-based\nspatial-temporal models. The study provides some guidance on robust model\ndesign and training: Transformer-based model performs better than CNN-based\nmodels on corruption robustness; the generalization ability of spatial-temporal\nmodels implies robustness against temporal corruptions; model corruption\nrobustness (especially robustness in the temporal domain) enhances with\ncomputational cost and model capacity, which may contradict the current trend\nof improving the computational efficiency of models. Moreover, we find the\nrobustness intervention for image-related tasks (e.g., training models with\nnoise) may not work for spatial-temporal models.\n

Paper

References (50)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC