Asking without Telling: Exploring Latent Ontologies in Contextual Representations

The success of pretrained contextual encoders, such as ELMo and BERT, has\nbrought a great deal of interest in what these models learn: do they, without\nexplicit supervision, learn to encode meaningful notions of linguistic\nstructure? If so, how is this structure encoded? To investigate this, we\nintroduce latent subclass learning (LSL): a modification to existing\nclassifier-based probing methods that induces a latent categorization (or\nontology) of the probe's inputs. Without access to fine-grained gold labels,\nLSL extracts emergent structure from input representations in an interpretable\nand quantifiable form. In experiments, we find strong evidence of familiar\ncategories, such as a notion of personhood in ELMo, as well as novel\nontological distinctions, such as a preference for fine-grained semantic roles\non core arguments. Our results provide unique new evidence of emergent\nstructure in pretrained encoders, including departures from existing\nannotations which are inaccessible to earlier methods.\n

Paper

References (84)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC