A Data Bootstrapping Recipe for Low Resource Multilingual Relation Classification

Relation classification (sometimes called 'extraction') requires trustworthy\ndatasets for fine-tuning large language models, as well as for evaluation. Data\ncollection is challenging for Indian languages, because they are syntactically\nand morphologically diverse, as well as different from resource-rich languages\nlike English. Despite recent interest in deep generative models for Indian\nlanguages, relation classification is still not well served by public data\nsets. In response, we present IndoRE, a dataset with 21K entity and relation\ntagged gold sentences in three Indian languages, plus English. We start with a\nmultilingual BERT (mBERT) based system that captures entity span positions and\ntype information and provides competitive monolingual relation classification.\nUsing this system, we explore and compare transfer mechanisms between\nlanguages. In particular, we study the accuracy efficiency tradeoff between\nexpensive gold instances vs. translated and aligned 'silver' instances. We\nrelease the dataset for future research.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC