Naamah: A Large Scale Synthetic Sanskrit NER Corpus via DBpedia Seeding and LLM Generation

The digitisation of classical Sanskrit literature is impeded by a scarcity of annotated resources, particularly for Named Entity Recognition. While recent methodologies utilise generic Large Language Models (LLMs) for data augmentation, these approaches remain prone to error and often lack the reasoning depth required for classical grammar. In this work, we introduce Naamah, a high quality silver standard Sanskrit NER dataset comprising 102,942 sentences. We propose a methodology that combines entity extraction from DBpedia with the generative capabilities of a 24B parameter hybrid reasoning model to create grammatically natural and synthetically diverse training data. We utilize this dataset to benchmark two transformer architectures: the massive multilingual XLM RoBERTa and the parameter efficient IndicBERTv2.

Paper

References (9)

03Sanskrit segmentation with LSTM networks2020 · Proceedings of the 12th Language Resources and Evaluation Conference
04DAGA: Data augmentation with a generation approach for low-resource tagging tasks2020 · Pro-ceedings of the 2020 Conference on Empirical Methods in Natural Language Processing
05Design and analysis of a Sanskrit sandhi splitter2016 · Proceedings of the 26th International Conference on Computational Linguistics
06DCS - the Digital Corpus of Sanskrit2010 · Linguistics, Archaeology and the Human Past, Occasional Paper 9 , Kyoto, Japan
08DBpedia mining strategy: We detail a methodology for extracting diverse entity seeds from DBpedia using structured queries, ensuring broad coverage of persons, locations, and organizations
09This serves as a strong multilingual baseline. It uses a large vocabulary of 250k tokens and is often the default choice for low resource languages. Howeverclassical Sanskrit

Similar papers

© 2026 NYSGPT2525 LLC