ATCSpeechNet: A multilingual end-to-end speech recognition framework for air traffic control systems

In this paper, a multilingual end-to-end framework, called as ATCSpeechNet,\nis proposed to tackle the issue of translating communication speech into\nhuman-readable text in air traffic control (ATC) systems. In the proposed\nframework, we focus on integrating the multilingual automatic speech\nrecognition (ASR) into one model, in which an end-to-end paradigm is developed\nto convert speech waveform into text directly, without any feature engineering\nor lexicon. In order to make up for the deficiency of the handcrafted feature\nengineering caused by ATC challenges, a speech representation learning (SRL)\nnetwork is proposed to capture robust and discriminative speech representations\nfrom the raw wave. The self-supervised training strategy is adopted to optimize\nthe SRL network from unlabeled data, and further to predict the speech\nfeatures, i.e., wave-to-feature. An end-to-end architecture is improved to\ncomplete the ASR task, in which a grapheme-based modeling unit is applied to\naddress the multilingual ASR issue. Facing the problem of small transcribed\nsamples in the ATC domain, an unsupervised approach with mask prediction is\napplied to pre-train the backbone network of the ASR model on unlabeled data by\na feature-to-feature process. Finally, by integrating the SRL with ASR, an\nend-to-end multilingual ASR framework is formulated in a supervised manner,\nwhich is able to translate the raw wave into text in one model, i.e.,\nwave-to-text. Experimental results on the ATCSpeech corpus demonstrate that the\nproposed approach achieves a high performance with a very small labeled corpus\nand less resource consumption, only 4.20% label error rate on the 58-hour\ntranscribed corpus. Compared to the baseline model, the proposed approach\nobtains over 100% relative performance improvement which can be further\nenhanced with the increasing of the size of the transcribed samples.\n

Paper

References (64)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC