Vyākarana: A Colorless Green Benchmark for Syntactic Evaluation in Indic Languages

While there has been significant progress towards developing NLU resources\nfor Indic languages, syntactic evaluation has been relatively less explored.\nUnlike English, Indic languages have rich morphosyntax, grammatical genders,\nfree linear word-order, and highly inflectional morphology. In this paper, we\nintroduce Vy\\=akarana: a benchmark of Colorless Green sentences in Indic\nlanguages for syntactic evaluation of multilingual language models. The\nbenchmark comprises four syntax-related tasks: PoS Tagging, Syntax Tree-depth\nPrediction, Grammatical Case Marking, and Subject-Verb Agreement. We use the\ndatasets from the evaluation tasks to probe five multilingual language models\nof varying architectures for syntax in Indic languages. Due to its prevalence,\nwe also include a code-switching setting in our experiments. Our results show\nthat the token-level and sentence-level representations from the Indic language\nmodels (IndicBERT and MuRIL) do not capture the syntax in Indic languages as\nefficiently as the other highly multilingual language models. Further, our\nlayer-wise probing experiments reveal that while mBERT, DistilmBERT, and XLM-R\nlocalize the syntax in middle layers, the Indic language models do not show\nsuch syntactic localization.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC