Low resource language dataset creation, curation and classification: Setswana and Sepedi -- Extended Abstract

The recent advances in Natural Language Processing have only been a boon for\nwell represented languages, negating research in lesser known global languages.\nThis is in part due to the availability of curated data and research resources.\nOne of the current challenges concerning low-resourced languages are clear\nguidelines on the collection, curation and preparation of datasets for\ndifferent use-cases. In this work, we take on the task of creating two datasets\nthat are focused on news headlines (i.e short text) for Setswana and Sepedi and\nthe creation of a news topic classification task from these datasets. In this\nstudy, we document our work, propose baselines for classification, and\ninvestigate an approach on data augmentation better suited to low-resourced\nlanguages in order to improve the performance of the classifiers.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC