Transformer-based audio self-supervised learning (SSL) models often use spectrogram time-frequency representations, vision-style Transformer architectures, and masked modeling objectives. They typically rely on convolutional patchification with temporal downsampling, which lowers the effective Nyquist frequency and introduces aliasing, whereas na’ve low-pass filtering can remove task-relevant high-frequency cues. We present AaSP, an aliasing-aware selfsupervised pre-training framework for audio spectrogram transformers. AaSP combines an aliasing-aware patch representation, teacher-student masked modeling, a cross-attention predictor, and multi-mask contrastive regularization to learn representations that better integrate features from alias-prone modulation bands while remaining stable across masked views. Its patch-embedding component, Aliasingaware Patch Embedding (AaPE), augments standard patch tokens with features from alias-prone modulation bands using a band-limited complex sinusoidal kernel with a two-sided exponential window. The kernel’s frequency and decay parameters are estimated from the input, enabling adaptive subband analysis whose outputs are fused with standard patch tokens. These components encourage consistency across masked views, which appears important for realizing the empirical benefit of the aliasing-aware patch representation. We pre-train on AudioSet and evaluate learned representations via fine-tuning and linear evaluation on downstream benchmarks spanning general acoustic/environmental, speech-related, and music-related audio recognition tasks. Under fine-tuning, the full AaSP framework achieves state-of-the-art results on AS-20 K, ESC-50, and NSynth among compared self-supervised baselines, while remaining competitive on other benchmarks. Linear evaluation shows a similar trend, with clear gains on several benchmarks, including US8K and NSynth. Overall, these results suggest that AaSP learns representations that are more stable under aliasingsensitive temporal perturbations and competitive for downstream transfer.
Paper
References (54)
Scroll for more · 38 remaining