Although BERT is widely used by the NLP community, little is known about its\ninner workings. Several attempts have been made to shed light on certain\naspects of BERT, often with contradicting conclusions. A much raised concern\nfocuses on BERT's over-parameterization and under-utilization issues. To this\nend, we propose o novel approach to fine-tune BERT in a structured manner.\nSpecifically, we focus on Large Scale Multilabel Text Classification (LMTC)\nwhere documents are assigned with one or more labels from a large predefined\nset of hierarchically organized labels. Our approach guides specific BERT\nlayers to predict labels from specific hierarchy levels. Experimenting with two\nLMTC datasets we show that this structured fine-tuning approach not only yields\nbetter classification results but also leads to better parameter utilization.\n
Paper
References (30)
Scroll for more · 18 remaining