How Familiar Does That Sound? Cross-Lingual Representational Similarity Analysis of Acoustic Word Embeddings
How do neural networks "perceive" speech sounds from unknown languages? Does\nthe typological similarity between the model's training language (L1) and an\nunknown language (L2) have an impact on the model representations of L2 speech\nsignals? To answer these questions, we present a novel experimental design\nbased on representational similarity analysis (RSA) to analyze acoustic word\nembeddings (AWEs) -- vector representations of variable-duration spoken-word\nsegments. First, we train monolingual AWE models on seven Indo-European\nlanguages with various degrees of typological similarity. We then employ RSA to\nquantify the cross-lingual similarity by simulating native and non-native\nspoken-word processing using AWEs. Our experiments show that typological\nsimilarity indeed affects the representational similarity of the models in our\nstudy. We further discuss the implications of our work on modeling speech\nprocessing and language similarity with neural networks.\n