Improving Speech Intelligibility using Deep Learning

Speech Intelligibility is of prime importance in places of mass gathering such as railway stations and airports as it is imperative for the speech to be understandable for the benefit and safety of the people. Often, the desired speech is corrupted by external noise which leads to discrepancies in the way it is perceived by the target audience. Therefore, the purpose of a speech enhancement algorithm is to identify and eliminate the noise. Recent advancements in speech processing and machine learning have resulted in various methods to do this. This research aims to improve speech intelligibility using Deep Neural Network (DNN) based speech enhancement system. Optimum Soft Mask (OSM) is used as a training target for the DNN system. The performance of the system is evaluated using Short Time Objective Intelligibility (STOI) metric. It is found that the proposed system provided good enhancement of noisy speech without compromising on intelligibility. In addition, it is found that usage of optimum soft mask instead of an ideal binary mask (IBM) provided additional improvement in speech intelligibility.

Paper

Full text

PDF

Improving Speech Intelligibility using Deep Learning

Semantic Scholar · Computer Science · 2019

Abstract

Speech Intelligibility is of prime importance in places of mass gathering such as railway stations and airports as it is imperative for the speech to be understandable for the benefit and safety of the people. Often, the desired speech is corrupted by external noise which leads to discrepancies in the way it is perceived by the target audience. Therefore, the purpose of a speech enhancement algorithm is to identify and eliminate the noise. Recent advancements in speech processing and machine learning have resulted in various methods to do this. This research aims to improve speech intelligibility using Deep Neural Network (DNN) based speech enhancement system. Optimum Soft Mask (OSM) is used as a training target for the DNN system. The performance of the system is evaluated using Short Time Objective Intelligibility (STOI) metric. It is found that the proposed system provided good enhancement of noisy speech without compromising on intelligibility. In addition, it is found that usage of optimum soft mask instead of an ideal binary mask (IBM) provided additional improvement in speech intelligibility.

Similar papers

© 2026 NYSGPT2525 LLC