Language Through a Prism: A Spectral Approach for Multiscale Language Representations

Language exhibits structure at different scales, ranging from subwords to\nwords, sentences, paragraphs, and documents. To what extent do deep models\ncapture information at these scales, and can we force them to better capture\nstructure across this hierarchy? We approach this question by focusing on\nindividual neurons, analyzing the behavior of their activations at different\ntimescales. We show that signal processing provides a natural framework for\nseparating structure across scales, enabling us to 1) disentangle\nscale-specific information in existing embeddings and 2) train models to learn\nmore about particular scales. Concretely, we apply spectral filters to the\nactivations of a neuron across an input, producing filtered embeddings that\nperform well on part of speech tagging (word-level), dialog speech acts\nclassification (utterance-level), or topic classification (document-level),\nwhile performing poorly on the other tasks. We also present a prism layer for\ntraining models, which uses spectral filters to constrain different neurons to\nmodel structure at different scales. Our proposed BERT + Prism model can better\npredict masked tokens using long-range context and produces multiscale\nrepresentations that perform better at utterance- and document-level tasks. Our\nmethods are general and readily applicable to other domains besides language,\nsuch as images, audio, and video.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC