Classification of language units (sounds, phonemes, lexemes) is an urgent task of computer linguistics. Its effective solution will allow, for example, to automate of the painstaking and time-consuming work of transcribing speech signals, which is necessary when creating speech corpora. Existing approaches to solving this problem mostly come from the field of automated speech recognition and are characterized by either extremely high requirements for the hardware component of the corresponding information system (with local implementation), or a low level of information security and saturated traffic (with network implementation). We also note that the a priori tendency of such systems to take into account the results of predicting the appearance of language units in speech signals in the process of classifying the first ones becomes a drawback when transcribing speech, for which a sufficiently developed universal background model is not available. In the thesis, a method of classification of language units is proposed, based on the Markov interpretation of parametrized cepstral patterns of the short-term representation of speech signals. The described method formalizes both the computationally efficient process of classifying language units based on the stationary distribution of the hidden Markov model of speech, and the training process of such a model, formulated with an orientation to the rational use of memory. Testing of the proposed method of classifying language units in the balanced metric of qualitative indicators showed its significant advantage over the classical approach in conditions where the number of speakers is relatively small and the size of the training sample is limited compared to the size of the test sample. Also, testing showed that the proposed method outperforms the classical method in terms of time spent on training and classification by at least two orders of magnitude.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex