Codeswitching is a widely observed phenomenon among bilingual speakers. By combining subword information enriched word vectors with linear-chain Conditional Random Field, we develop a supervised machine learning model that identifies languages in a English-Spanish codeswitched tweets. Our computational method achieves a tweet-level weighted F1 of 0.83 and a token-level accuracy of 0.949 without using any exter-nal resource. The result demonstrates that named entity recognition remains a challenge in codeswitched texts and warrants further work.
Paper
Full text
Codeswitching language identification using Subword Information Enriched Word Vectors
Semantic Scholar · Computer Science · 2016
Abstract
Codeswitching is a widely observed phenomenon among bilingual speakers. By combining subword information enriched word vectors with linear-chain Conditional Random Field, we develop a supervised machine learning model that identifies languages in a English-Spanish codeswitched tweets. Our computational method achieves a tweet-level weighted F1 of 0.83 and a token-level accuracy of 0.949 without using any exter-nal resource. The result demonstrates that named entity recognition remains a challenge in codeswitched texts and warrants further work.
References (19)
Scroll for more · 7 remaining