Spanish Biomedical Crawled Corpus: A Large, Diverse Dataset for Spanish Biomedical Language Models
We introduce CoWeSe (the Corpus Web Salud Espa\\~nol), the largest Spanish\nbiomedical corpus to date, consisting of 4.5GB (about 750M tokens) of clean\nplain text. CoWeSe is the result of a massive crawler on 3000 Spanish domains\nexecuted in 2020. The corpus is openly available and already preprocessed.\nCoWeSe is an important resource for biomedical and health NLP in Spanish and\nhas already been employed to train domain-specific language models and to\nproduce word embbedings. We released the CoWeSe corpus under a Creative Commons\nAttribution 4.0 International license, both in Zenodo\n(\\url{https://zenodo.org/record/4561971\\#.YTI5SnVKiEA}).\n