A common framework of identifying bird species from audio recordings involves detecting bird song segments, which will be subsequently input to a classifier. In-field recordings are contaminated with various environmental noise. For such recordings, supervised segmentation has been observed to outperform unsupervised energy-based approaches. Prior supervised segmentation work considers only pixel-level predictions and ignores the supervision provided at the segment-level. We propose a hierarchical approach that learns to isolate bird song syllables based on both pixel-level and segment-level information. Experimental results suggest that our method outperforms an existing supervised method that learns only from pixel-level supervision.
Paper
Full text
Supervised hierarchical segmentation for bird song recording
Semantic Scholar · Computer Science · 2015
Abstract
A common framework of identifying bird species from audio recordings involves detecting bird song segments, which will be subsequently input to a classifier. In-field recordings are contaminated with various environmental noise. For such recordings, supervised segmentation has been observed to outperform unsupervised energy-based approaches. Prior supervised segmentation work considers only pixel-level predictions and ignores the supervision provided at the segment-level. We propose a hierarchical approach that learns to isolate bird song syllables based on both pixel-level and segment-level information. Experimental results suggest that our method outperforms an existing supervised method that learns only from pixel-level supervision.