Fine Tuning Methods for Low-resource Languages

The rise of Large Language Models has not been inclusive of all cultures. The models are mostly trained on English texts and culture which makes them underperform in other languages and cultural contexts. By developing a generalizable method for preparing culturally relevant datasets and post-training the Gemma 2 model, this project aimed to increase the performance of Gemma 2 for an underrepresented language and showcase how others can do the same to unlock the power of Generative AI in their country and preserve their cultural heritage.

Paper

References (23)

08The dominance of english-trained ai systems and the risk of digital language death: Preserving linguistic diversity for cultural identity and knowledge transmission2020 · Publications of the National Research Council of Italy
09LitteraturbankenLitteraturbanken: Free swedish literature in digital format
10Spoken language task-specific tuning with gemma, 2024Google AI
11Gensim fasttext implementation:/
122025 overviewLt4all

Scroll for more · 11 remaining

Similar papers

© 2026 NYSGPT2525 LLC