PTT5: Pretraining and validating the T5 model on Brazilian Portuguese data

In natural language processing (NLP), there is a need for more resources in\nPortuguese, since much of the data used in the state-of-the-art research is in\nother languages. In this paper, we pretrain a T5 model on the BrWac corpus, an\nextensive collection of web pages in Portuguese, and evaluate its performance\nagainst other Portuguese pretrained models and multilingual models on three\ndifferent tasks. We show that our Portuguese pretrained models have\nsignificantly better performance over the original T5 models. Moreover, we\ndemonstrate the positive impact of using a Portuguese vocabulary. Our code and\nmodels are available at https://github.com/unicamp-dl/PTT5.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC