Convolutional Recurrent Neural Networks for the Classification of Cetacean Bioacoustic Patterns

In this paper we focus on the development of a convolutional recurrent neural network (CRNN) to categorize biosignals collected in the Hellenic Trench, generated by two cetacean species, sperm whales (Physeter macrocephalus) and striped dolphins (Stenella coeruleoalba). We convert audio signals into mel-spectrograms and forward the input into a deep residual network (ResNet), designed to capture spectral patterns. Next, ResNet’s output is reshaped into a time-distributed layer and fed into recurrent network variants, Long Short-Term Memory (LSTMs) or Gated Recurrent Units (GRUs), able to recognize long-term time dependencies on extracted features. The hybrid network perfectly classifies audio signals into three categories (dolphins, sperm whales, ambient noise) while it also exhibits high learning ability on recognising intraclass representations of overlapping acoustic patterns (clicks vs whistles and clicks, both emitted by dolphins). The proposed scheme outperforms traditional Machine Learning (ML) techniques, baseline ResNet and LSTM architectures or their deep parallel combinations.

Paper

Full text

PDF

Convolutional Recurrent Neural Networks for the Classification of Cetacean Bioacoustic Patterns

Semantic Scholar · Computer Science · 2023

Abstract

In this paper we focus on the development of a convolutional recurrent neural network (CRNN) to categorize biosignals collected in the Hellenic Trench, generated by two cetacean species, sperm whales (Physeter macrocephalus) and striped dolphins (Stenella coeruleoalba). We convert audio signals into mel-spectrograms and forward the input into a deep residual network (ResNet), designed to capture spectral patterns. Next, ResNet’s output is reshaped into a time-distributed layer and fed into recurrent network variants, Long Short-Term Memory (LSTMs) or Gated Recurrent Units (GRUs), able to recognize long-term time dependencies on extracted features. The hybrid network perfectly classifies audio signals into three categories (dolphins, sperm whales, ambient noise) while it also exhibits high learning ability on recognising intraclass representations of overlapping acoustic patterns (clicks vs whistles and clicks, both emitted by dolphins). The proposed scheme outperforms traditional Machine Learning (ML) techniques, baseline ResNet and LSTM architectures or their deep parallel combinations.

Similar papers

© 2026 NYSGPT2525 LLC