Not All Models Localize Linguistic Knowledge in the Same Place: A Layer-wise Probing on BERToids' Representations

Most of the recent works on probing representations have focused on BERT,\nwith the presumption that the findings might be similar to the other models. In\nthis work, we extend the probing studies to two other models in the family,\nnamely ELECTRA and XLNet, showing that variations in the pre-training\nobjectives or architectural choices can result in different behaviors in\nencoding linguistic information in the representations. Most notably, we\nobserve that ELECTRA tends to encode linguistic knowledge in the deeper layers,\nwhereas XLNet instead concentrates that in the earlier layers. Also, the former\nmodel undergoes a slight change during fine-tuning, whereas the latter\nexperiences significant adjustments. Moreover, we show that drawing conclusions\nbased on the weight mixing evaluation strategy -- which is widely used in the\ncontext of layer-wise probing -- can be misleading given the norm disparity of\nthe representations across different layers. Instead, we adopt an alternative\ninformation-theoretic probing with minimum description length, which has\nrecently been proven to provide more reliable and informative results.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC