Boosting Transformers for Job Expression Extraction and Classification in a Low-Resource Setting

In this paper, we explore possible improvements of transformer models in a\nlow-resource setting. In particular, we present our approaches to tackle the\nfirst two of three subtasks of the MEDDOPROF competition, i.e., the extraction\nand classification of job expressions in Spanish clinical texts. As neither\nlanguage nor domain experts, we experiment with the multilingual XLM-R\ntransformer model and tackle these low-resource information extraction tasks as\nsequence-labeling problems. We explore domain- and language-adaptive\npretraining, transfer learning and strategic datasplits to boost the\ntransformer model. Our results show strong improvements using these methods by\nup to 5.3 F1 points compared to a fine-tuned XLM-R model. Our best models\nachieve 83.2 and 79.3 F1 for the first two tasks, respectively.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC