Does BERT Learn as Humans Perceive? Understanding Linguistic Styles through Lexica

People convey their intention and attitude through linguistic styles of the\ntext that they write. In this study, we investigate lexicon usages across\nstyles throughout two lenses: human perception and machine word importance,\nsince words differ in the strength of the stylistic cues that they provide. To\ncollect labels of human perception, we curate a new dataset, Hummingbird, on\ntop of benchmarking style datasets. We have crowd workers highlight the\nrepresentative words in the text that makes them think the text has the\nfollowing styles: politeness, sentiment, offensiveness, and five emotion types.\nWe then compare these human word labels with word importance derived from a\npopular fine-tuned style classifier like BERT. Our results show that the BERT\noften finds content words not relevant to the target style as important words\nused in style prediction, but humans do not perceive the same way even though\nfor some styles (e.g., positive sentiment and joy) human- and\nmachine-identified words share significant overlap for some styles.\n

Paper

References (20)

Scroll for more · 8 remaining

Similar papers

© 2026 NYSGPT2525 LLC