We present hinglishNorm -- a human annotated corpus of Hindi-English\ncode-mixed sentences for text normalization task. Each sentence in the corpus\nis aligned to its corresponding human annotated normalized form. To the best of\nour knowledge, there is no corpus of Hindi-English code-mixed sentences for\ntext normalization task that is publicly available. Our work is the first\nattempt in this direction. The corpus contains 13494 parallel segments.\nFurther, we present baseline normalization results on this corpus. We obtain a\nWord Error Rate (WER) of 15.55, BiLingual Evaluation Understudy (BLEU) score of\n71.2, and Metric for Evaluation of Translation with Explicit ORdering (METEOR)\nscore of 0.50.\n
Paper
References (52)
Scroll for more · 38 remaining