Seeing Both the Forest and the Trees: Multi-head Attention for Joint Classification on Different Compositional Levels

In natural languages, words are used in association to construct sentences.\nIt is not words in isolation, but the appropriate combination of hierarchical\nstructures that conveys the meaning of the whole sentence. Neural networks can\ncapture expressive language features; however, insights into the link between\nwords and sentences are difficult to acquire automatically. In this work, we\ndesign a deep neural network architecture that explicitly wires lower and\nhigher linguistic components; we then evaluate its ability to perform the same\ntask at different hierarchical levels. Settling on broad text classification\ntasks, we show that our model, MHAL, learns to simultaneously solve them at\ndifferent levels of granularity by fluidly transferring knowledge between\nhierarchies. Using a multi-head attention mechanism to tie the representations\nbetween single words and full sentences, MHAL systematically outperforms\nequivalent models that are not incentivized towards developing compositional\nrepresentations. Moreover, we demonstrate that, with the proposed architecture,\nthe sentence information flows naturally to individual words, allowing the\nmodel to behave like a sequence labeller (which is a lower, word-level task)\neven without any word supervision, in a zero-shot fashion.\n

Paper

References (46)

Scroll for more · 34 remaining

Similar papers

© 2026 NYSGPT2525 LLC