hinglishNorm -- A Corpus of Hindi-English Code Mixed Sentences for Text Normalization

We present hinglishNorm -- a human annotated corpus of Hindi-English\ncode-mixed sentences for text normalization task. Each sentence in the corpus\nis aligned to its corresponding human annotated normalized form. To the best of\nour knowledge, there is no corpus of Hindi-English code-mixed sentences for\ntext normalization task that is publicly available. Our work is the first\nattempt in this direction. The corpus contains 13494 parallel segments.\nFurther, we present baseline normalization results on this corpus. We obtain a\nWord Error Rate (WER) of 15.55, BiLingual Evaluation Understudy (BLEU) score of\n71.2, and Metric for Evaluation of Translation with Explicit ORdering (METEOR)\nscore of 0.50.\n

Paper

References (52)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC