We describe the CMU submission for the 2014 shared task on language identification in code-switched data. We participated in all four language pairs: Spanish‐English, Mandarin‐English, Nepali‐English, and Modern Standard Arabic‐Arabic dialects. After describing our CRF-based baseline system, we discuss three extensions for learning from unlabeled data: semi-supervised learning, word embeddings, and word lists.
Paper
Full text
The CMU Submission for the Shared Task on Language Identification in Code-Switched Data
Semantic Scholar · Computer Science · 2014
Abstract
We describe the CMU submission for the 2014 shared task on language identification in code-switched data. We participated in all four language pairs: Spanish‐English, Mandarin‐English, Nepali‐English, and Modern Standard Arabic‐Arabic dialects. After describing our CRF-based baseline system, we discuss three extensions for learning from unlabeled data: semi-supervised learning, word embeddings, and word lists.
References (17)
Scroll for more · 5 remaining