Codeswitching language identification using Subword Information Enriched Word Vectors

Codeswitching is a widely observed phenomenon among bilingual speakers. By combining subword information enriched word vectors with linear-chain Conditional Random Field, we develop a supervised machine learning model that identifies languages in a English-Spanish codeswitched tweets. Our computational method achieves a tweet-level weighted F1 of 0.83 and a token-level accuracy of 0.949 without using any exter-nal resource. The result demonstrates that named entity recognition remains a challenge in codeswitched texts and warrants further work.

Paper

Full text

PDF

Codeswitching language identification using Subword Information Enriched Word Vectors

Semantic Scholar · Computer Science · 2016

Abstract

Codeswitching is a widely observed phenomenon among bilingual speakers. By combining subword information enriched word vectors with linear-chain Conditional Random Field, we develop a supervised machine learning model that identifies languages in a English-Spanish codeswitched tweets. Our computational method achieves a tweet-level weighted F1 of 0.83 and a token-level accuracy of 0.949 without using any exter-nal resource. The result demonstrates that named entity recognition remains a challenge in codeswitched texts and warrants further work.

References (19)

Scroll for more · 7 remaining

Similar papers

© 2026 NYSGPT2525 LLC